{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2010/07/19/scripts-logs-as-htmlrdfa-mix-free-text.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/tb3sc-ggw42",
      "url": "https://chem-bla-ics.linkedchemistry.info/2010/07/19/scripts-logs-as-htmlrdfa-mix-free-text.html",
      "title": "Script logs as HTML+RDFa: mix free text reporting with CSV",
      "content_html": "<p><a href=\"http://blogs.talis.com/nodalities/author/richard-wallis/\">Richard</a> (<a href=\"http://www.talis.com/\">Talis</a>) wrote up a\n<a href=\"http://blogs.talis.com/nodalities/2010/07/the-data-publishing-three-step.php\">three-step tutorial</a> on how to publish\nyour data. I think I would be more than happy if scientists reached step 1. Related, Ola asked me a while ago if I\nwas interested in using the computing facilities of <a href=\"http://www.uppmax.uu.se/\">UPPMAX</a>, and I was. But until this\nweekend I did not have the time or energy to give it a spin. If you are puzzled how the heck I see those two items\nrelated, read on :)</p>\n\n<p>Two days later, today, I ran my first analysis. Still a test run, but using the <a href=\"http://cdk.sf.net/\">CDK</a> to perceive\natom types on the first 2.5 GB of <a href=\"http://pubchem.ncbi.nlm.nih.gov/\">PubChem</a> data. The full data set is now 80 GB,\nand I will start doing this analysis today. You might remember this already two years ago (see\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2008/05/03/wicked-chemistry-and-unit-testing.html\">Wicked chemistry and unit testing <i class=\"fa-solid fa-recycle fa-xs\"></i></a>)\nfor a small subset, but only now have the power to analyze all compounds. The UPPMAX system I work on has 348, each\nwith 8 cores. Each core has 3 GB of memory, but I am using the\n<a href=\"http://pele.farmbio.uu.se/nightly-1.2.3/cdk-javadoc-1.2.4/org/openscience/cdk/io/iterator/IteratingPCCompoundXMLReader.html\">IteratingPCCompoundXMLReader</a>\nclass anyway. Analyzing the 2.5 GB of data was done using 50 nodes, and finished in about a minute. Nice :)</p>\n\n<p>Now, this first run dumped the results as a plain text file, looking like:</p>\n\n<div class=\"language-plaintext highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code>CID 200234: Ti  1\nCID 200235: Ti 1\nCID 200237: Sb 1 Sb 2\nCID 200365: S 1\nCID 200761: Hg 1\nCID 201374: Ce 1 Ce 2\nCID 201395: As 1 \n</code></pre></div></div>\n\n<p>Simple and effective.</p>\n\n<p>Or? And this is where the two items outlined in the first paragraph meet. No, this is not useful. Since the output\nis from an analysis of PubChem, I’m sure you already figured out that the first two columns indicate the compound\nbeing analyzed. You might also work out that then the elements are given for which the atom type perception failed.\nYou may even figure out that the number is likely to be the index in the connection table representation of the\nmolecule. Right?</p>\n\n<p>But what about machine readability? I could, of course, write the output as CSV, but then I would loose my ability\nto write the report in human readable format. And moreover, the list of failing atom types does not have a fixed\nlength, as you can see in the example lines given earlier.</p>\n\n<p>Now, this is where RDF comes in. If I create my output as HTML+RDFa, I can do fancy stuff. My results page could\nlink directly to PubChem, so that I can inspect the actual compound. Though I could do that even with merely HTML.\nBut with <a href=\"http://www.w3.org/TR/xhtml-rdfa-primer/\">RDFa</a>, I can actually make my free text log output machine\nreadable. I can accurately annotate what bits are informative:</p>\n\n<div class=\"language-xml highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code><span class=\"nt\">&lt;div</span> <span class=\"na\">about=</span><span class=\"s\">\"#200234\"</span> <span class=\"na\">typeof=</span><span class=\"s\">\"um:Compound\"</span><span class=\"nt\">&gt;</span>CID\n  <span class=\"nt\">&lt;span</span> <span class=\"na\">property=</span><span class=\"s\">\"um:cid\"</span> <span class=\"na\">datatype=</span><span class=\"s\">\"xsd:integer\"</span><span class=\"nt\">&gt;</span>200234<span class=\"nt\">&lt;/span&gt;</span>:\n  <span class=\"nt\">&lt;span</span> <span class=\"na\">rel=</span><span class=\"s\">'um:hasProblem'</span><span class=\"nt\">&gt;</span>\n    <span class=\"nt\">&lt;span</span> <span class=\"na\">about=</span><span class=\"s\">'#error0'</span> <span class=\"na\">typeof=</span><span class=\"s\">'um:Problem'</span><span class=\"nt\">&gt;</span>\n      <span class=\"nt\">&lt;span</span> <span class=\"na\">property=</span><span class=\"s\">'um:hasElement'</span><span class=\"nt\">&gt;</span>Ti<span class=\"nt\">&lt;/span&gt;</span>\n      <span class=\"nt\">&lt;span</span> <span class=\"na\">property=</span><span class=\"s\">'um:hasIndex'</span> <span class=\"na\">datatype=</span><span class=\"s\">'xsd:integer'</span><span class=\"nt\">&gt;</span>1<span class=\"nt\">&lt;/span&gt;</span>\n    <span class=\"nt\">&lt;/span&gt;</span>\n  <span class=\"nt\">&lt;/span&gt;</span>\n<span class=\"nt\">&lt;/div&gt;</span>\n</code></pre></div></div>\n\n<p>The file is not backed up by an OWL ontology, but where possible one would do that. Reuse of ontologies is a good\nthing (e.g. use a service like <a href=\"http://schemapedia.com/\">Schemapedia</a>).</p>\n\n<p>Now, I can easily open up this file in a web browser (follow <a href=\"http://rdf.farmbio.uu.se/uppmax-cdk/results.html\">this link</a>)\nand get the same view as above. But I can also import the file directly into Bioclipse (see\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/01/28/semantic-web-features-in-bioclipse-22.html\">Semantic Web features in Bioclipse 2.2 <i class=\"fa-solid fa-recycle fa-xs\"></i></a>),\nor in any other tool that supports RDFa. I can then use SPARQL to do some first analysis, for example, with:</p>\n\n<div class=\"language-sparql highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code><span class=\"k\">PREFIX</span><span class=\"w\"> </span><span class=\"nn\">um</span><span class=\"o\">:</span><span class=\"w\"> </span><span class=\"nn\">&lt;http://egonw.github.com/uppmax&gt;</span><span class=\"w\">\n\n</span><span class=\"k\">SELECT</span><span class=\"w\"> </span><span class=\"nv\">?elem</span><span class=\"w\"> </span><span class=\"p\">(</span><span class=\"nb\">count</span><span class=\"p\">(</span><span class=\"o\">*</span><span class=\"p\">)</span><span class=\"w\"> </span><span class=\"k\">AS</span><span class=\"w\"> </span><span class=\"nv\">?count</span><span class=\"p\">)</span><span class=\"w\"> </span><span class=\"k\">WHERE</span><span class=\"w\"> </span><span class=\"p\">{</span><span class=\"w\">\n  </span><span class=\"nv\">?compound</span><span class=\"w\"> </span><span class=\"nn\">um</span><span class=\"o\">:</span><span class=\"ss\">cid</span><span class=\"w\"> </span><span class=\"nv\">?cid</span><span class=\"p\">;</span><span class=\"w\">\n     </span><span class=\"nn\">um</span><span class=\"o\">:</span><span class=\"ss\">hasProblem</span><span class=\"w\"> </span><span class=\"nv\">?problem</span><span class=\"w\"> </span><span class=\"p\">.</span><span class=\"w\">\n  </span><span class=\"nv\">?problem</span><span class=\"w\"> </span><span class=\"nn\">um</span><span class=\"o\">:</span><span class=\"ss\">hasElement</span><span class=\"w\"> </span><span class=\"nv\">?elem</span><span class=\"w\"> </span><span class=\"p\">.</span><span class=\"w\">\n</span><span class=\"p\">}</span><span class=\"w\"> </span><span class=\"k\">GROUP</span><span class=\"w\"> </span><span class=\"k\">BY</span><span class=\"w\"> </span><span class=\"nv\">?elem</span><span class=\"w\"> </span><span class=\"k\">ORDER</span><span class=\"w\"> </span><span class=\"k\">BY</span><span class=\"w\"> </span><span class=\"nv\">?elem</span><span class=\"w\">\n</span></code></pre></div></div>\n\n<p>Combine that with the <a href=\"http://rdfadev.sourceforge.net/\">RDFaDev</a> tool I wrote about last week (see\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/07/16/rdfadev-htmlrdfa-development-with.html\">RDFaDev: HTML+RDFa development with FireFox <i class=\"fa-solid fa-recycle fa-xs\"></i></a>).\nNow you should get some feeling of the advantages of using Open Standards: I can do some initial analysis of the results,\njust right there in the web browser you have open anyway:</p>\n\n<p><img src=\"/assets/images/rdfaLogfiles.png\" alt=\"\" /></p>\n\n<p>Therefore, next time you ask your data analyst to perform some calculation, insist that he sends you HTML+RDFa log files with\nresults. Better, ask him to put it online, and you immediately reach\n<a href=\"http://blogs.talis.com/nodalities/2010/07/the-data-publishing-three-step.php\">Step 3</a>\nin the analysis by David.</p>",
      "summary": "Richard (Talis) wrote up a three-step tutorial on how to publish your data. I think I would be more than happy if scientists reached step 1. Related, Ola asked me a while ago if I was interested in using the computing facilities of UPPMAX, and I was. But until this weekend I did not have the time or energy to give it a spin. If you are puzzled how the heck I see those two items related, read on :)",
      "image": "https://chem-bla-ics.linkedchemistry.info/assets/images/rdfaLogfiles.png",
      "date_published": "2010-07-19T00:00:00+00:00",
      "date_modified": "2026-09-27T00:00:00+00:00",
      "tags": ["html","rdf","sparql","rdfa"],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
