{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2008/08/03/end-of-theory-data-deluge-makes.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/j4d10-0jt05",
      "url": "https://chem-bla-ics.linkedchemistry.info/2008/08/03/end-of-theory-data-deluge-makes.html",
      "title": "&quot;The End of Theory: The Data Deluge Makes the Scientific Method Obsolete&quot;",
      "content_html": "<p>The thought triggering editorial <a href=\"http://www.wired.com/science/discoveries/magazine/16-07/pb_theory\">“The End of Theory: The Data Deluge Makes the Scientific Method Obsolete”</a>\nby <a href=\"http://www.wired.com/services/feedback/letterstoeditor\">Chris Anderson</a> can’t have escaped your attention. I was shocked when I read the title\nand the comments made on the blogosphere and on <a href=\"http://friendfeed.com/\">FriendFeed</a>.</p>\n\n<p>How can he say that?! There is no analysis of data anymore?!? Don’t we need to understand why X correlated with Y?!? Etc etc.</p>\n\n<p>So, when I read <a href=\"http://miningdrugs.blogspot.com/2008/08/data-models-or-both.html\">yet another comment</a>, by my respected opensource\nchemoinformatician <a href=\"http://miningdrugs.blogspot.com/\">Joerg</a>, I just had to read the piece myself. Joerg disagrees with the statement\nfrom Chris’ editorial that</p>\n\n<blockquote>\n  <p>[c]orrelation supersedes causation, and science can advance even without coherent models, unified theories, or really any\nmechanistic explanation at all.</p>\n</blockquote>\n\n<p>At first, I would agree with Joerg. It’s nonsense; any QSAR modeler can explain in details the dangers of overfitting, extrapolation,\netc, etc. Not to mention that basically zero mathematical modeling methods can create a statistical signification non-zero regression model with less than 50-100 chemical structures (chemical diversity dependent, etc).</p>\n\n<p>Ok, back to the editorial. There are some arguments on Google, tons of data. Number of incoming links as measure of page importance (brilliant choice, but actually a model, IMHO, which Chris seems to step over). Tons of data. Oh, mentioned that already.</p>\n\n<p>Mmmmm… but wait. Tons of data? The editorial actually refers to <a href=\"http://en.wikipedia.org/wiki/Petabyte\">petabytes</a>:\n<em>Petabytes are stored in the cloud</em>. (Whatever the cloud is… just another buzzword,\n<a href=\"http://news.slashdot.org/article.pl?sid=08/08/02/2224217&amp;from=rss\">trademarketed too</a>, it seems).</p>\n\n<h2 id=\"eureka-chris-is-right-joerg-is-wrong\">Eureka! Chris is right, Joerg is wrong!</h2>\n\n<p>Yes! Then it hit me, Chris is actually correct in his statement, and I was wrong (and Joerg too). If we move away from 50-100 molecules\nin our QSAR training, but use 10k of chemically alike molecules, then our modeling approaches (if capable of handling the matrices)\nwould have a much, much smaller chance for overfitting, extrapolation (there is much, much more interpolation now), etc. The chances\nof getting random correlation become insignificant! Actually, Chris is making the argument QSAR modelists have been making for decades:\nwe do not know the mode of action in detail, as we can make, given enough training data, a reasonable regression model to predict the\naction! Joerg and I have been making the same argument as Chris in our PhD theses! We do not need theory; our QSAR regressions make\ntheory obsolete! (Well, surely, we’d still prefer the theory behind the action, but we lack the measuring techniques to see what\nactually is happening. Joerg, still agreeing with you, so to say ;)</p>\n\n<p>Except for one thing. Joerg and I suggested ‘enough’ molecules are required for statistical sound regression. Chris, on the other hand,\neven makes the point that regression is no longer needed at all at the petabyte scene: we just look up what is happening. Does this hold\nfor chemistry? For QSAR? Petabyte data equals about, say 10kB data per structure, maybe less if we use InChI and neglect conformer info,\n100.000.000.000 structures. About 5000 times <a href=\"http://chemspider.com/\">ChemSpider</a>, if not miscounting the zeros (we don’t care about a\nten-fold at this scale anymore). Maybe, maybe not. Maybe chemical space is too diverse for that, considering a petabyte of chemical\nstructures is enormously insignificant to the full drugable space (was about 10⁶⁰, not?)</p>\n\n<p>But not at all? This lookup approach is actually commonly used in chemoinformatics! Even at a way-below-pentybyte scale:\nHOSE-code-based NMR prediction is a nice example of this! We do not theorize on the chemical carbon NMR shift, we just look it up!</p>\n\n<p>Certainly worth reading, this <a href=\"http://www.wired.com/\">Wired</a> editorial!</p>\n\n<p>PS. One last remark on the title… I’d say the the <em>scientific method</em> is more than just making theories… I feel a bit left\nout as data analyst… :( I guess the title should have said ‘one of the Scientific Methods’…</p>",
      "summary": "The thought triggering editorial “The End of Theory: The Data Deluge Makes the Scientific Method Obsolete” by Chris Anderson can’t have escaped your attention. I was shocked when I read the title and the comments made on the blogosphere and on FriendFeed.",
      
      "date_published": "2008-08-03T00:00:00+00:00",
      "date_modified": "2008-08-03T00:00:00+00:00",
      "tags": ["cheminf"],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
