{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2010/12/26/oscar-training-data-models-etc.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/pa72q-ykk64",
      "url": "https://chem-bla-ics.linkedchemistry.info/2010/12/26/oscar-training-data-models-etc.html",
      "title": "Oscar: training data, models, etc",
      "content_html": "<p><a href=\"https://sourceforge.net/projects/oscar3-chem/\">Oscar</a> uses a Maximum Entropy Markov Model (MEMM) based on <a href=\"http://en.wikipedia.org/wiki/N-gram\">n-grams</a>.\nPeter Corbett has written this up (doi:<a href=\"https://doi.org/10.1186/1471-2105-9-S11-S4\">10.1186/1471-2105-9-S11-S4</a>). So, it basically is statistics\nonce more. If you really want a proper bioinformatics education, so do your PhD at a (proteo)chemometrics department.</p>\n\n<p>N-grams are word parts of n characters. For example, the trigrams of <a href=\"http://en.wikipedia.org/wiki/Acetic_acid\">acetic acid</a>\ninclude <code class=\"language-plaintext highlighter-rouge\">ace</code>, <code class=\"language-plaintext highlighter-rouge\">cid</code>, <code class=\"language-plaintext highlighter-rouge\">tic</code>, <code class=\"language-plaintext highlighter-rouge\">eti</code>, and <code class=\"language-plaintext highlighter-rouge\">aci</code>. N-grams of length four include acid, etic, and acet. The MEMM assigns weights to\nthese n-grams, and based on that decided if something is in deed a <em>named entity</em> (in Oscar terminology). For example,\nconsider the <code class=\"language-plaintext highlighter-rouge\">acet</code> n-gram: acetone should be matched, but the n-gram <code class=\"language-plaintext highlighter-rouge\">facet</code> not.</p>\n\n<p>Put this in perspective in the ongoing refactoring of the Oscar software. We are changing normalization (e.g. converting\nall unicode hyphen alternatives into one specific hyphen), updating the tokenizer (e.g. changing the list of\nnon-sentence-endings like <em>Prof.</em>). It is clear this changes the n-grams typical for chemical-like things. Worse,\nthe weights are tuned towards to know n-grams, and statistical models are generally a bit overtrained for the\ndata, or, at least, specific for it.</p>\n\n<p>Now, if the distribution of n-grams changes, the weights in the model need to be updated too, to not degrade\nthe model performance. So, Oscar is useless if we cannot retrain its MEMM component after a refactoring. If\nthat would be impossible, we would have effectively created an <em>intellectual monopoly</em>.</p>\n\n<p>Thus, what the Oscar project needs, is one or more free sets of annotated literature, which can be used to\ntrain new MEMM models. The SciBorg corpus was used to train the current Oscar3 and Oscar4 models. This data\n(copyright <a href=\"http://rsc.org/\">RSC</a>) will very likely be available under a <a href=\"http://creativecommons.org/licenses/\">Creative Commons</a>\nlicense (RSC++), but may have the NC clause, which would not be good for developing a business model around\nthe opensource Oscar (such as providing a high-performance web service via a subscription service). I have\nrecently written up <a href=\"http://chem-bla-ics.blogspot.com/2010/12/re-why-i-and-you-should-avoid-nc.html\">the problems the NC clause introduces</a>,\nand some <a href=\"http://chem-bla-ics.blogspot.com/2010/12/blog-post.html\">examples of commercial Open Source cheminformatics projects</a>.</p>\n\n<p>We need not focus only on this SciBorg data, however. In fact, we will need multiple models anyway. For\nexample, the SciBorg papers (42 if not mistaken) are around a particular kind of literature. So, it\nintroduces the risk of using it to analyse papers out of the application domain. Furthermore, I am very\ninterested (and others indicated so too) to use Oscar for other languages. Surely, English is the major\nlanguage, but there are many use cases for Oscar when useful for other languages.</p>\n\n<p>Therefore, for what we need in the Oscar project, is a registry of training (/test) data, annotated itself\nwith metadata around how that data was created (what quality assurance, what kind of named entity types,\nhow many domain experts were involved, etc), test results for those data sets, etc. My time on the Oscar\nproject is almost over, and I have no clue when I will be able to invest the same amount of time into the\nproject as I did in the past three months. But the creation of this registry is clear step that must be\ntaken in the Oscar4 development.</p>\n\n<h4>References</h4>\n<div class=\"csl-bib-body\">\n    <div class=\"csl-entry\">Corbett, P., &#38; Copestake, A. (2008). Cascaded classifiers for confidence-based chemical named entity recognition. <i>BMC Bioinformatics</i>, <i>9</i>(S11). https://doi.org/10.1186/1471-2105-9-s11-s4 <a href=\"https://doi.org/10.1186/1471-2105-9-S11-S4\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/1471-2105-9-S11-S4\">Scholia</a></div>\n  </div>",
      "summary": "Oscar uses a Maximum Entropy Markov Model (MEMM) based on n-grams. Peter Corbett has written this up (doi:10.1186/1471-2105-9-S11-S4). So, it basically is statistics once more. If you really want a proper bioinformatics education, so do your PhD at a (proteo)chemometrics department.",
      
      "date_published": "2010-12-26T00:00:00+00:00",
      "date_modified": "2010-12-26T00:00:00+00:00",
      "tags": ["oscar","textmining"],
      "_references": [
        
          
          
            { "url": "https://doi.org/10.1186/1471-2105-9-S11-S4", "doi": "10.1186/1471-2105-9-S11-S4"
             }
            
          
        ],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
