{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2007/07/13/inter-and-extrapolation-nmr-shift.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/tbd0q-67564",
      "url": "https://chem-bla-ics.linkedchemistry.info/2007/07/13/inter-and-extrapolation-nmr-shift.html",
      "title": "Inter- and Extrapolation: the NMR shift prediction debate",
      "content_html": "<p>Chemical blogspace has seen a lengthy discussion on <a href=\"https://chem-bla-ics.linkedchemistry.info/2007/06/19/quality-of-chemical-database.html\">the quality of a few NMR shift prediction programs <i class=\"fa-solid fa-recycle fa-xs\"></i></a>,\nand Ryan wanted to make <a href=\"http://acdlabs.typepad.com/my_weblog/2007/07/final-note-on-t.html\">a final statement</a>. Down his blog item\nhe had this quote from Jeff, discussing the use of the <a href=\"http://www.nmrshiftdb.org/\">NMRShiftDB</a> as external test set:</p>\n\n<blockquote>\n  <p>“Of course customers are really interested in how accurately a prediction program can predict THEIR molecules - not a collection of external data such as NMRShiftDB.”</p>\n</blockquote>\n\n<p>I’m sure none of us knows what weird chemistry people are doing; we will never know what the overlap of the NMRShiftDB test\nset with the customer data set is. The quote suggests it is low, but we simply do not know.</p>\n\n<h2 id=\"interpolation-and-extrapolation\">Interpolation and Extrapolation</h2>\n\n<p>The accuracy of prediction models is very difficult to grasp, and one can only estimate it; using a test set.\nIf few data is available, one may opt for using the training set as test set too, and gives an estimate if the\nmodeling method is able to predict at all. However, the outcome of this exercise is the worst possible estimate\nyou can make. So, when possible you use an independent test set, which does not contain any molecules that were\npresent in the training set. (Actually, one could even suggest that this must happen on a shift level, but that\ngives problems with HOSE-code based prediction.)</p>\n\n<p>Now, what Ryan stresses in his <a href=\"http://acdlabs.typepad.com/my_weblog/2007/07/final-note-on-t.html\">latest blog item</a>\nis that prediction test results for the various available methods does not explicitly state the amount of overlap\nbetween the training and test set, one cannot draw any conclusions. Agreed. I would, however, like to tune this\neven a bit further, after reading the stupid quote (of course, taking out of context). What Jeff probably aimed\nat, is that the prediction accuracy is only meaningful to a customer if there is considerable between the customers\ndata set and the test set, which is what the model makers do not know.</p>\n\n<p>And the overlap actually goes beyond the overlap in terms of molecular identity. It is really the overlap in terms\nof molecular substructures that matters: a database with alkanes but no phenyl rings will more accurately predict\nother alkanes not present in the training set (interpolation), but will not accurately predict compounds with\nphenyl rings (extrapolation). What the customer needs is that his personal data set does not require extrapolation.\nThat is what matters.</p>\n\n<p>It is interesting to realize, however, that the NMRShiftDB allows you to upload your molecules, or alternatively,\nyou download the software (it’s open source) and the data (it’s open data) if you don’t want to send your molecules\nover the internet, and the NMRShiftDB software will automatically take into account your own data set.</p>\n\n<p>Thus, if you are working on a series of related molecules, you can extend the NMRShiftDB data set with already\nelucidated structures, reducing the prediction error for your yet related unknowns derivatives. It is that easy\nto include prior/expert knowledge in the NMRShiftDB. I believe the ACD/Labs software allows this too, so the\nquote is really meaningless. Not correct, not wrong, simply says nothing.</p>\n\n<h2 id=\"open-data-open-source-open-standards\">Open Data, Open Source, Open Standards</h2>\n\n<p>Now, the various releases of the ACD/Labs software show a simple, understandable trend that increasing the number\nof data you use for the training set, reduces the prediction error. That’s because of various reasons I will not\ngo into in this item. The ACD/Labs NMR databases are expensive, because they have to manually extract and validate\nthe data from literature (see <a href=\"http://acdlabs.typepad.com/my_weblog/2007/06/the_purgatory_d.html\">The Purgatory Database</a>);\nso, during my PhD I only bought the CNMR and HNMR prediction packages. (Off topic: two weeks after I received my\ncopies of the software, ACD/Labs released a new version, which they kindly sent me a copy of too. Common in\nopensource, but much appreciated at that time. Cheers, <a href=\"http://www.acdlabs.com/\">ACD/Labs</a>!)</p>\n\n<p>The ACD/Labs databases are likely expensive because of various reasons. And this is where the ODOSOS concept of the\n<a href=\"http://www.blueobelisk.org/\">Blue Obelisk</a> comes in. <strong>Open Data</strong>: if publishers would not copyright their data,\nNMR databases would be much cheaper to set up (see <a href=\"https://blogs.ch.cam.ac.uk/pmr/2007/07/12/do-authors-want-to-give-publishers-a-monopoly-over-their-data/\">this thread in Peter’s blog <i class=\"fa-solid fa-recycle fa-xs\"></i></a>);\nassuming ACD/Labs has to pay publishers for actually setting up their database. <strong>Open Source</strong>: the various Blue\nObelisk projects provide the <a href=\"https://chem-bla-ics.linkedchemistry.info/2006/09/08/chemical-archeology-oscar3-to.html\">tools to automatically create a purgatory NMR database <i class=\"fa-solid fa-recycle fa-xs\"></i></a>;\nno humans needed for that any more. <strong>Open Standards</strong>: the data from the NMRShiftDB can be downloaded in various\nformats, among which CMLSpect. Being able to easily read the data, made it possible that we actually have this\ndiscussion. Sure, the open data part of the NMRShiftDB is crucial too! But the database could have used an obscure,\nbinary, undocumented, with many software tweaks and special cases, <code class=\"language-plaintext highlighter-rouge\">.doc</code>-like format, which no one could support.</p>\n\n<p>Clearly, ODOSOS gives all, even proprietary, NMR prediction tools a boost, and I am very happy to see that happen.\nIt is the point that we, the Blue Obelisk Movement, are trying to make for some time now.</p>",
      "summary": "Chemical blogspace has seen a lengthy discussion on the quality of a few NMR shift prediction programs , and Ryan wanted to make a final statement. Down his blog item he had this quote from Jeff, discussing the use of the NMRShiftDB as external test set:",
      
      "date_published": "2007-07-13T00:00:00+00:00",
      "date_modified": "2025-07-30T00:00:00+00:00",
      "tags": ["nmr","nmrshiftdb"],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
