{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2005/11/08/when-to-stop-including-qsar-model.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/hxb0r-66s49",
      "url": "https://chem-bla-ics.linkedchemistry.info/2005/11/08/when-to-stop-including-qsar-model.html",
      "title": "When to stop including QSAR model variables...",
      "content_html": "<p>Yesterday I reviewed an article which published a QSPR model which looked something like:</p>\n\n\\[y = 151 + 50p1 - 12p2 - 0.006p3\\]\n\n<p>with quite OK prediction results (R=0.9880). But I was not quite comfortable with the coefficient for the \\(p3\\) variable.\nThe article did not calculate significances for the coefficients, so it was not obvious from the article wether is was useful\nto include them. I then looked at the range for <code class=\"language-plaintext highlighter-rouge\">p3</code>, which was 110-150; so, the maximal influence this variable can have is\n\\(150*0.006 = 0.9\\). Now, the experimental values given in the article were rounded to integers, indicating that the maximal\neffect of the <code class=\"language-plaintext highlighter-rouge\">p3</code> variable is smaller than the experimental error! It’s even worse when you consider the difference between the\nmin and max value (40), then the influence would even be smaller (assuming that most model methods would put the mean temperature\neffect in the offset, 151 in this case).</p>\n\n<p>Today, I reread an article with a similar issue. The model was something like:</p>\n\n\\[y = -0.81 + 0.03*p1 + 0.009*p2\\]\n\n<p>Here, \\(max(p2)-min(p2)\\) is a smaller than 100, so the maximal effect of the variable would be in the order 0.9, which is of\nthe same order of the root mean square error of prediction (RMSEP) for this model. Indeed, the article already states that the\ncoefficient is only significant at the 95% level, and not at the 99% level. But, without having calculated the RMSEP for a model\nwithout the p4 variable, I would guess that leaving it out would give equally good prediction results.</p>\n\n<p>Concluding, I would say the the <code class=\"language-plaintext highlighter-rouge\">p2</code> variable does not include relevant information.</p>\n\n<p>Do you think it is reasonable to include the <code class=\"language-plaintext highlighter-rouge\">p2</code> variable in the second model?</p>",
      "summary": "Yesterday I reviewed an article which published a QSPR model which looked something like:",
      
      "date_published": "2005-11-08T00:00:00+00:00",
      "date_modified": "2005-11-08T00:00:00+00:00",
      "tags": ["cheminf","qsar"],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
