{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2025/04/27/one-million-iupac-names-2-the-100-thousand-milestone.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/dycsw-qeq51",
      "url": "https://chem-bla-ics.linkedchemistry.info/2025/04/27/one-million-iupac-names-2-the-100-thousand-milestone.html",
      "title": "One Million IUPAC names #2: the 100 thousand milestone",
      "content_html": "<p>Two and a half month into the <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html\">One Million IUPAC Names</a>\nproject, we passed <a href=\"https://github.com/BlueObelisk/iupac-names/releases/tag/milestone-100k\">the third milestone</a>,\nthe one for 100 thousand IUPAC names (doi:<a href=\"https://doi.org/10.5281/zenodo.15266459\">10.5281/zenodo.15266459</a>).\nTime for an update.</p>\n\n<p>This milestone release took a bit longer. Going from 50 to 100 thousand is a bigger step than from 10 to 50\nthousand, but the open access chemistry literature was already done by then. Basically, I ran out of open access\nchemistry publications. The scripts are now finding names in all (open access) literature, and the number of\nnew names per articles is a lot lower. Still about 1 in every twenty to 30 articles. But the diversity in names\nis not really going down, which is important.</p>\n\n<p>The first few weeks, I used the Google Colab to run a Jupyter notebook, initial created by\n<a href=\"https://cpm.lumc.nl/research/bioinformatics-224/magnus-palmblad-5\">Magnus</a>, but having to process more articles\nto get a reasonable number of new IUPAC names required longer and longer jobs, and then Google Colab\nis not really fit (well, the free version anyway). So, I started using a local script. That turned out\nto be able to handle up to 20 thousand articles in one go and runs at least twice as fast. Moreover, I can\nrun three of them in parallel.</p>\n\n<p>And that had impact. With each commit around 1000 new IUPAC names, the number of commits went up remarkably\nlast week:</p>\n\n<p><img src=\"/assets/images/iupac-names-commits.png\" alt=\"\" /></p>\n\n<p>At the current speed, I think we’ll make it to 150k soon and I added a new milestone for 200k, which sounds\ndoable in the next three week. That also means that 1M extracted IUPAC names from literature has become\na reasonable goal. And we can start thinking about the 2, 5, 10, 50 and 100 million IUPAC names. Those are,\nat the current speed, rather unlikely to reach from the open access literature anytime soon. That brings\nus to the question, what will. Well, I have some ideas.</p>\n\n<h3 id=\"idea-1-name-variations\">Idea 1: name variations</h3>\n\n<p>First, I am figuring out some ways to make variants of names (no, not based on hyphens and spaces; that’s too easy),\nbut actual variations of the chemical structures. For example, I could exhaustively replace “methoxy” with “ethoxy”,\nand iterate the halogens and acyl chain lengts. I have little doubt that I can grow the list with this approach\neasily a 5-fold, maybe even a 10-fold.</p>\n\n<h3 id=\"idea-2-hallucination\">Idea 2: hallucination</h3>\n\n<p>Another idea is that I could use tools that can generate IUPAC names for a limited set of compounds.\nI once wrote code for alkanes myself and if I can find that, I may be able to generate additional names.\nBut perhaps more realistic is that I train a deep learning model and have it generate names for all compounds in\nWikidata (~1.5 million) or PubChem (&gt;100 million). STOUT needed 81 million compounds\n(doi:<a href=\"https://doi.org/10.1186/s13321-021-00512-4\">10.1186/s13321-021-00512-4</a>), but I don’t need a good model;\nI just need a model that comes up with new, valid names. Hallucinated names, but valid.</p>\n\n<p>While the list of valid names grows, I can retrain the deep-learned model and repeat. As long as the diversity\nremains high enough, one could hypothesize that the deep learning will learn new tricks. And then,\nthat should be a near infinite source of additional names.</p>\n\n<h3 id=\"idea-3-semi-closed-access-literature\">Idea 3: (semi-)closed access literature</h3>\n\n<p>Also, I haven’t touched closed access articles yet. This is all based on the collection of full texts\nin <a href=\"https://europepmc.org/\">Europe PMC</a>. For example, I could start with the green open access article\nin (Dutch) university repositories, particularly those with large chemistry departments. PDF to text\ntools are mature enough that this will provide a new source. Oh, and perhaps PhD thesis, which are now\nalso increasingly archived in university repository under open access. And that reminds me of a Dutch\nproject two decades ago doing exactly that. I wish I remembered the name.</p>\n\n<h3 id=\"idea-4-alternatives-to-oscar4-and-europe-pmc\">Idea 4: alternatives to Oscar4 and Europe PMC</h3>\n\n<p>So, the first round of named entity recognition was with Europe PMC itself, as explained in\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html\">the first post</a>. The move\nto Oscar4 helped a lot. But there exist many other chemical NER tools, like\n(doi:<a href=\"https://doi.org/10.1093/bioinformatics/btn181\">10.1093/bioinformatics/btn181</a>. And those may\nfind an additional number of names, even with just the literature I already covered.</p>\n\n<p>Well, you get the idea.</p>\n\n<h2 id=\"iccs-poster-rejected\">ICCS poster rejected</h2>\n\n<p>Unfortunately, the <a href=\"https://iccs-nl.org/\">ICCS poster</a> abstract did not make the cut. The score was high enough,\nbut they received many abstracts and had to make a selection (of course, I am part of the ICCS organization,\nand have more details of how it came about). I really like the project, and eager to write up a paper around\nit.</p>\n\n\n<h4>References</h4>\n<div class=\"csl-bib-body\">\n    <div class=\"csl-entry\">Klinger, R., Kolářik, C., Fluck, J., Hofmann-Apitius, M., &#38; Friedrich, C. M. (2008). Detection of IUPAC and IUPAC-like chemical names. <i>Bioinformatics</i>, <i>24</i>(13), i268–i276. https://doi.org/10.1093/bioinformatics/btn181 <b>[cito:citesAsPotentialSolution]</b> <a href=\"https://doi.org/10.1093/bioinformatics/btn181\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1093/bioinformatics/btn181\">Scholia</a></div>\n    <div class=\"csl-entry\">Rajan, K., Zielesny, A., &#38; Steinbeck, C. (2021). STOUT: SMILES to IUPAC names using neural machine translation. <i>Journal of Cheminformatics</i>, <i>13</i>(1). https://doi.org/10.1186/s13321-021-00512-4 <b>[cito:citesForInformation]</b> <a href=\"https://doi.org/10.1186/s13321-021-00512-4\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/s13321-021-00512-4\">Scholia</a></div>\n  </div>",
      "summary": "Two and a half month into the One Million IUPAC Names project, we passed the third milestone, the one for 100 thousand IUPAC names (doi:10.5281/zenodo.15266459). Time for an update.",
      "image": "https://chem-bla-ics.linkedchemistry.info/assets/images/iupac-names-commits.png",
      "date_published": "2025-04-27T00:00:00+00:00",
      "date_modified": "2025-06-09T00:00:00+00:00",
      "tags": ["iupac","textmining","oscar","europepmc"],
      "_references": [
        
          
          
            { "url": "https://doi.org/10.1186/s13321-021-00512-4", "doi": "10.1186/s13321-021-00512-4"
            , "cito":
              
              
                [ 
                  "citesForInformation"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1093/bioinformatics/btn181", "doi": "10.1093/bioinformatics/btn181"
            , "cito":
              
              
                [ 
                  "citesAsPotentialSolution"
                  
                 ]
              
             }
            
          
        ],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
