{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/tjkf2-k1608",
      "url": "https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html",
      "title": "One Million IUPAC names",
      "content_html": "<p>Names of chemicals are part of the human user experience when browsing a chemical database. And literature too,\nof course. Chemical names are also not easy to use, and what a chemical name means is not always clear.\nThis is why the <a href=\"https://en.wikipedia.org/wiki/International_Union_of_Pure_and_Applied_Chemistry\">IUPAC</a>\nstarted a standardizing nomenclature in chemistry, the <em>IUPAC names</em>. Each IUPAC name uniquely defines\nthe chemical structure it defines. For example, <em>methane</em> is the IUPAC name for the chemical CH<sub>4</sub>.</p>\n\n<p>So, when propagating chemical structures from the <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html\">Beilstein Bioschemas feed</a>,\nI was looking for names, IUPAC or not, ideally the name used in the article. When I asked about this,\nthe question came up if they could autogenerate IUPAC names, for which\n<a href=\"https://doi.org/10.1038/s41598-021-94082-y\">various</a>\n<a href=\"https://doi.org/10.1186/s13321-021-00535-x\">new</a>\n<a href=\"https://doi.org/10.1186/s13321-021-00512-4\">tools</a>\n<a href=\"https://doi.org/10.1186/s13321-024-00941-x\">exist</a>\n(I think I am missing one from an American team, but cannot find the reference),\nalong with multiple established commerical tools.\nBecause the IUPAC nomenclature is a long list of naming rules, priorities, etc, a rule-based\nalgorithm is logical, but newer methods take a deep-learning approach.</p>\n\n<p>Back to the chemical annotation of chemistry literature. This is of obvious interest: you want\nto know where we can read more about a certain chemical. We need the chemical structures in\na database for that, linked to the articles. This is, of course, one of the original studies\nof <em>cheminformatics</em>. And when authors of the chemical literature do not provide this routinely\n(<a href=\"https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html\">this post</a>\nshows a few exceptions, but it is still all too rare). And then manual and automated curation\nis needed, e.g. done by <a href=\"https://en.wikipedia.org/wiki/Chemical_Abstracts_Service\">Chemical Abstracts</a>.</p>\n\n<p>Third, <a href=\"https://wikidata.org/\">Wikidata</a> has <a href=\"https://scholia.toolforge.org/chemical/\">about 1.4 million</a>\nchemical compounds and many names. A <a href=\"https://www.wikidata.org/wiki/Wikidata:Property_proposal/Pending#IUPAC_name\">property propoal for IUPAC names</a>\nhas been long pending, but once accepted in one form or another, will require IUPAC names too.</p>\n\n<h2 id=\"one-million-iupac-names\">One million IUPAC names</h2>\n\n<p>Thus, the idea came up, can we create a set of 1 million unique IUPAC names found in literature?\nI asked on the <a href=\"https://elixir-europe.org/\">ELIXIR Europe</a> slack channel if <a href=\"https://europepmc.org/\">Europe PMC</a>\nhad such a dataset (doi:<a href=\"https://doi.org/10.1093/nar/gkad1085\">10.1093/nar/gkad1085</a>). I knew they had been adding chemical\n<a href=\"https://scholia.toolforge.org/topic/Q403574\">named-entity recognition</a> (NER) results in\n<a href=\"https://europepmc.org/Annotations\">their annotation API</a>. I learned they used <a href=\"https://www.ebi.ac.uk/chebi/\">ChEBI</a>.\nMelanie Vollmar and Summer Rosonovski or Europe PMC gave useful information and support.\n<a href=\"https://cpm.lumc.nl/research/bioinformatics-224/magnus-palmblad-5\">Magnus Palmblad</a> also replied\nand provided Python code to use the Europe PMC API to fetch names it returns and see if those\nare IUPAC names. Well, that’s easy. We have <a href=\"https://opsin.ch.cam.ac.uk/\">OPSIN</a> for that\n(see doi:<a href=\"https://doi.org/10.1021/ci100384d\">10.1021/ci100384d</a>).</p>\n\n<p>Unfortunately, the Europe PMC NER results are not ideal for IUPAC names. Just scanning\nsome 5, 6 organic chemistry journals returned some 8 thousand IUPAC names in open access\narticles. But it quickly started to be too limited: each set of articles returned\nincreasingly few new names. The reason is simple: the NER is too <em>greedy</em> and as a\nresult, does not easily recognize longer IUPAC names. It is too happy with a substring\nof the IUPAC name. For example, when it encounters the IUPAC name <em>5-Bromo-1H-indole-3-carboxylic acid</em>,\nit settles for <em>indole-3-carboxylic acid</em>:</p>\n\n<p><img src=\"/assets/images/greedy.png\" alt=\"\" /></p>\n\n<h2 id=\"open-source-chemistry-analysis-routines\">Open-Source Chemistry Analysis Routines</h2>\n\n<p>During my PhD, in 2003, when I worked a few months with Prof. <a href=\"https://scholia.toolforge.org/author/Q908710\">Peter Murray-Rust</a> (University of Cambridge)\nand Prof. Janet Thornthon (EMBL-EBI), I learned about the research by <a href=\"https://scholia.toolforge.org/author/Q28946549\">Sam Adams</a>\n(doi:<a href=\"https://doi.org/10.1039/B411699M\">10.1039/B411699M</a>), <a href=\"https://scholia.toolforge.org/author/Q133040220\">Joe Townsend</a>\n(doi:<a href=\"https://doi.org/10.1039/B411033A\">10.1039/B411033A</a>), and <a href=\"https://scholia.toolforge.org/author/Q90318722\">Peter Corbett</a>\n(doi:<a href=\"https://doi.org/10.1007/11875741_11\">10.1007/11875741_11</a>). One of the tools that used\nthis research was (is) <a href=\"https://scholia.toolforge.org/topic/Q133037490\">OSCAR</a>,\nshort for <em>Open-Source Chemistry Analysis Routines</em> (see <a href=\"https://blogs.ch.cam.ac.uk/pmr/2009/05/16/opsin-and-oscar-chemical-language-processing/\">this detailed write up by Peter MR</a>).\nLater, in 2010 I visted Peter again, as postdoc, in Cambridge, and then\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/10/15/working-on-oscar-for-three-months.html\">worked on the OSCAR project</a> too.\nAnd while OSCAR did a lot more, the integration of <a href=\"https://chem-bla-ics.linkedchemistry.info/2010/12/26/oscar-training-data-models-etc.html\">Corbett’s NER research</a>\nmade OSCAR the obvious follow-up step in finding IUPAC names in literature.</p>\n\n<p>And because <a href=\"https://chem-bla-ics.linkedchemistry.info/2011/09/27/almost-year-ago-i-started-position-with.html\">OSCAR4 had been integrated into Bioclipse</a>\n(doi:<a href=\"https://doi.org/10.1186/1758-2946-3-41\">10.1186/1758-2946-3-41</a>) and I had this ported to Bacting already\n(doi:<a href=\"https://doi.org/10.21105/joss.02558\">10.21105/joss.02558</a>), using this was trivial.\nThe use of Europe PMC is different now, however, and we are no longer using the Annotations API,\nbut just using it to find open access articles, and to get the full text in XML format.\nThat allows a simple XPath search on <code class=\"language-plaintext highlighter-rouge\">&lt;p&gt;</code> elements, pass the resulting string to OSCAR4,\nand the recognized names are checked with OPSIN.\nAnd with this approach, processing two of the five or six journals we earlier explored,\nwe find another 40+ thousand IUPAC names. Quite a success, I am tempted to say.</p>\n\n<h2 id=\"a-blue-obelisk-project\">A Blue Obelisk project</h2>\n\n<p>So, I started a new <a href=\"https://blueobelisk.github.io/\">Blue Obelisk</a> project,\n<a href=\"https://github.com/BlueObelisk/iupac-names\">iupac-names</a>, to collect 1M IUPAC names. For researchers\nto use, learn from, etc. Just IUPAC names. Not even the chemical structure, nor the link to the\narticles. The first is trivial to do with OPSIN, so the matching SMILES do not need to be stored.\nLinks to literature is tricky because of the aforementioned issues, and we only want to know\nwhich (partial) IUPAC names occur in literature. If you really want to know in which articles\nthat IUPAC name is found, you can simply do a search in Europe PMC.</p>\n\n<p>And because we only store IUPAC names, this are very basic facts (this is an IUPAC name, as defined\nby OPSIN being able to generate a SMILES for this structure) and that that string occurs in\nsome article) and we can share them as CCZero. We <a href=\"is:issue\" title=\"milestone release\">defined various milestones</a>,\nand I am happy that the first two have been reached within two weeks:</p>\n\n<ul>\n  <li><a href=\"https://github.com/BlueObelisk/iupac-names/releases/tag/milestone-10k\">Milestone 10k</a> (doi:<a href=\"https://doi.org/10.5281/zenodo.14965762\">10.5281/zenodo.14965762</a>)</li>\n  <li><a href=\"https://github.com/BlueObelisk/iupac-names/releases/tag/milestone-50k\">Milestone 50k</a> (doi:<a href=\"https://doi.org/10.5281/zenodo.14978557\">10.5281/zenodo.14978557</a>)</li>\n</ul>\n\n<p>This second milestone has 53848 unique names, but as literature goes, there are interesting\nvariations, some likely because of typesetting leading to spaces added and missing. If\nwe ignore spaces and hyphens, we have 50534 names left (hence the milestone). But IUPAC\nnames are also not fully unique, partly because of Unicode character variations and greek\nletter alternatives, and you may wonder how many different chemical structures this set\nreflects. While not perfect, the Standard InChI gives some lower limit, and we find 36528\nInChIKeys in this second milestone.</p>\n\n<p>Now, we need twenty times as much to reach the 1M IUPAC names, but given we have many, many\nmore open access articles to process. The bottleneck seems to be mostly our workflow.</p>\n\n<h3 id=\"can-you-contribute\">Can you contribute?</h3>\n\n<p>Yes, of course! This is an open science project. But please keep in mind the narrow focus of this\nproject: only IUPAC names which can be found in (open access) literature. This project doed not accept\nautogenerated names (PubChem would have given use many millions already), nor IUPAC names from existing\ndatabases. Ideally, you are able to show the code you use to extract/find those names in literature.</p>\n\n<h3 id=\"can-i-use-these-names\">Can I use these names?</h3>\n\n<p>First of all, this is what the CCZero license and open science nature of this project is about: reuse.\nWe love to hear how you are using these names, tho, and we encourage you to write up how you\nare using them. You can use <a href=\"https://datacite.org/\">DataCite</a> to cite the release you used,\nand citing this blog post by DOI is also possible.</p>\n\n<h3 id=\"does-it-support-my-language-too\">Does it support my language too?</h3>\n\n<p>No, at this moment it only support IUPAC names in English. Dutch, French, Spanish, or Chinese\nIUPAC names are valid, but currently not supported. See also\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/12/30/text-mining-chemistry-from-dutch-or.html\">this post</a>.</p>\n\n<h3 id=\"will-there-be-a-publication\">Will there be a publication?</h3>\n\n<p>Magnus and I intend so. We already submitted an abstract to the <a href=\"https://iccs-nl.org/\">International Conference on Chemical Structures</a>,\nwhich has <a href=\"https://www.biomedcentral.com/collections/ICCS25\">a Collection in the Journal of Cheminformatics</a>.\nIf the abstract gets accepted, of course, we can submit there. Otherwise, we will look for another venue,\nlikely <a href=\"https://en.wikipedia.org/wiki/Diamond_open_access\">diamond open access</a>.</p>\n\n<h3 id=\"where-is-your-script\">Where is your script?</h3>\n\n<p>Ah, fair point. We did not decide on the final license yet. I have used two scripts based on the template\nby Magnus. As soon as we have finalized the license, we will make those available.</p>\n\n<h4>References</h4>\n<div class=\"csl-bib-body\">\n    <div class=\"csl-entry\">Adams, S. E., Goodman, J. M., Kidd, R. J., McNaught, A. D., Murray-Rust, P., Norton, F. R., Townsend, J. A., &#38; Waudby, C. A. (2004). Experimental data checker: better information for organic chemists. <i>Organic &#38;amp;  Biomolecular Chemistry</i>, <i>2</i>(21), 3067. https://doi.org/10.1039/b411699m <a href=\"https://doi.org/10.1039/B411699M\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1039/B411699M\">Scholia</a></div>\n    <div class=\"csl-entry\">Corbett, P., &#38; Murray-Rust, P. (2006). High-Throughput Identification of Chemistry in Life Science Texts. In <i>Lecture Notes in Computer Science</i> (pp. 107–118). Springer Berlin Heidelberg. https://doi.org/10.1007/11875741_11 <a href=\"https://doi.org/10.1007/11875741_11\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1007/11875741_11\">Scholia</a></div>\n    <div class=\"csl-entry\">Egon Willighagen. (2025). <i>BlueObelisk/iupac-names: Milestone 10k</i> (Version milestone-10k) [Dataset]. Zenodo. https://doi.org/10.5281/ZENODO.14965762 <b>[cito:citesAsEvidence]</b> <a href=\"https://doi.org/10.5281/zenodo.14965762\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.5281/zenodo.14965762\">Scholia</a></div>\n    <div class=\"csl-entry\">Egon Willighagen. (2025). <i>BlueObelisk/iupac-names: Milestone 50k</i> (Version milestone-50k) [Dataset]. Zenodo. https://doi.org/10.5281/ZENODO.14978557 <b>[cito:citesAsEvidence]</b> <a href=\"https://doi.org/10.5281/zenodo.14978557\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.5281/zenodo.14978557\">Scholia</a></div>\n    <div class=\"csl-entry\">Handsel, J., Matthews, B., Knight, N. J., &#38; Coles, S. J. (2021). Translating the InChI: adapting neural machine translation to predict IUPAC names from a chemical identifier. <i>Journal of Cheminformatics</i>, <i>13</i>(1). https://doi.org/10.1186/s13321-021-00535-x <a href=\"https://doi.org/10.1186/s13321-021-00535-x\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/s13321-021-00535-x\">Scholia</a></div>\n    <div class=\"csl-entry\">Jessop, D. M., Adams, S. E., Willighagen, E. L., Hawizy, L., &#38; Murray-Rust, P. (2011). OSCAR4: a flexible architecture for chemical text-mining. <i>Journal of Cheminformatics</i>, <i>3</i>(1). https://doi.org/10.1186/1758-2946-3-41 <b>[cito:usesMethodIn]</b> <a href=\"https://doi.org/10.1186/1758-2946-3-41\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/1758-2946-3-41\">Scholia</a></div>\n    <div class=\"csl-entry\">Krasnov, L., Khokhlov, I., Fedorov, M. V., &#38; Sosnin, S. (2021). Transformer-based artificial neural networks for the conversion between chemical notations. <i>Scientific Reports</i>, <i>11</i>(1). https://doi.org/10.1038/s41598-021-94082-y <a href=\"https://doi.org/10.1038/s41598-021-94082-y\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1038/s41598-021-94082-y\">Scholia</a></div>\n    <div class=\"csl-entry\">Lowe, D. M., Corbett, P. T., Murray-Rust, P., &#38; Glen, R. C. (2011). Chemical Name to Structure: OPSIN, an Open Source Solution. <i>Journal of Chemical Information and Modeling</i>, <i>51</i>(3), 739–753. https://doi.org/10.1021/ci100384d <a href=\"https://doi.org/10.1021/ci100384d\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1021/ci100384d\">Scholia</a></div>\n    <div class=\"csl-entry\">Rajan, K., Zielesny, A., &#38; Steinbeck, C. (2021). STOUT: SMILES to IUPAC names using neural machine translation. <i>Journal of Cheminformatics</i>, <i>13</i>(1). https://doi.org/10.1186/s13321-021-00512-4 <a href=\"https://doi.org/10.1186/s13321-021-00512-4\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/s13321-021-00512-4\">Scholia</a></div>\n    <div class=\"csl-entry\">Rajan, K., Zielesny, A., &#38; Steinbeck, C. (2024). STOUT V2.0: SMILES to IUPAC name conversion using transformer models. <i>Journal of Cheminformatics</i>, <i>16</i>(1). https://doi.org/10.1186/s13321-024-00941-x <a href=\"https://doi.org/10.1186/s13321-024-00941-x\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/s13321-024-00941-x\">Scholia</a></div>\n    <div class=\"csl-entry\">Rosonovski, S., Levchenko, M., Bhatnagar, R., Chandrasekaran, U., Faulk, L., Hassan, I., Jeffryes, M., Mubashar, S. I., Nassar, M., Jayaprabha Palanisamy, M., Parkin, M., Poluru, J., Rogers, F., Saha, S., Selim, M., Shafique, Z., Ide-Smith, M., Stephenson, D., Tirunagari, S., … Harrison, M. (2023). Europe PMC in 2023. <i>Nucleic Acids Research</i>, <i>52</i>(D1), D1668–D1676. https://doi.org/10.1093/nar/gkad1085 <b>[cito:usesMethodIn]</b> <a href=\"https://doi.org/10.1093/nar/gkad1085\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1093/nar/gkad1085\">Scholia</a></div>\n    <div class=\"csl-entry\">Townsend, J. A., Adams, S. E., Waudby, C. A., de Souza, V. K., Goodman, J. M., &#38; Murray-Rust, P. (2004). Chemical documents: machine understanding and automated information extraction. <i>Organic &#38;amp;  Biomolecular Chemistry</i>, <i>2</i>(22), 3294. https://doi.org/10.1039/b411033a <a href=\"https://doi.org/10.1039/B411033A\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1039/B411033A\">Scholia</a></div>\n    <div class=\"csl-entry\">Willighagen, E. (2021). Bacting: a next generation, command line version of Bioclipse. <i>Journal of Open Source Software</i>, <i>6</i>(62), 2558. https://doi.org/10.21105/joss.02558 <b>[cito:usesMethodIn]</b> <a href=\"https://doi.org/10.21105/JOSS.02558\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.21105/JOSS.02558\">Scholia</a></div>\n  </div>",
      "summary": "Names of chemicals are part of the human user experience when browsing a chemical database. And literature too, of course. Chemical names are also not easy to use, and what a chemical name means is not always clear. This is why the IUPAC started a standardizing nomenclature in chemistry, the IUPAC names. Each IUPAC name uniquely defines the chemical structure it defines. For example, methane is the IUPAC name for the chemical CH4.",
      "image": "https://chem-bla-ics.linkedchemistry.info/assets/images/greedy.png",
      "date_published": "2025-03-08T00:00:00+00:00",
      "date_modified": "2025-03-12T00:00:00+00:00",
      "tags": ["iupac","cheminf","oscar","textmining","europepmc"],
      "_references": [
        
          
          
            { "url": "https://doi.org/10.1038/s41598-021-94082-y", "doi": "10.1038/s41598-021-94082-y"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1186/s13321-021-00512-4", "doi": "10.1186/s13321-021-00512-4"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1186/s13321-021-00535-x", "doi": "10.1186/s13321-021-00535-x"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1186/s13321-024-00941-x", "doi": "10.1186/s13321-024-00941-x"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1021/ci100384d", "doi": "10.1021/ci100384d"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1039/B411699M", "doi": "10.1039/B411699M"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1039/B411033A", "doi": "10.1039/B411033A"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1007/11875741_11", "doi": "10.1007/11875741_11"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1186/1758-2946-3-41", "doi": "10.1186/1758-2946-3-41"
            , "cito":
              
              
                [ 
                  "usesMethodIn"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.21105/JOSS.02558", "doi": "10.21105/JOSS.02558"
            , "cito":
              
              
                [ 
                  "usesMethodIn"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1093/nar/gkad1085", "doi": "10.1093/nar/gkad1085"
            , "cito":
              
              
                [ 
                  "usesMethodIn"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.5281/zenodo.14965762", "doi": "10.5281/zenodo.14965762"
            , "cito":
              
              
                [ 
                  "citesAsEvidence"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.5281/zenodo.14978557", "doi": "10.5281/zenodo.14978557"
            , "cito":
              
              
                [ 
                  "citesAsEvidence"
                  
                 ]
              
             }
            
          
        ],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
