{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2009/09/04/nmrshiftdb-enters-rdfopenmoleculesnet-2.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/nv925-tje87",
      "url": "https://chem-bla-ics.linkedchemistry.info/2009/09/04/nmrshiftdb-enters-rdfopenmoleculesnet-2.html",
      "title": "NMRShiftDB enters rdf.openmolecules.net #2: SPARQL end point with Virtuoso",
      "content_html": "<p>About 6 months ago I <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/03/18/nmrshiftdb-enters-rdfopenmoleculesnet.html\">reported <i class=\"fa-solid fa-recycle fa-xs\"></i></a> about my efforts to RDF-ize the data from the\n<a href=\"http://www.nmrshiftdb.org/\">NMRShiftDB</a>. Since then, time was consumed by many other things, but now that <a href=\"http://www.bioclipse.net/\">Bioclipse</a> can query\n<a href=\"http://en.wikipedia.org/wiki/SPARQL\">SPARQL</a> end points, that I want to contribute the triple set (it is <a href=\"http://www.gnu.org/copyleft/fdl.html\">GNU FDL</a>-licensed)\nto <a href=\"http://www.bio2rdf.org/\">Bio2RDF</a>, that a student started working in my group (now larger than just me :) on reasoning on life sciences data, and that I\nrecently contributed my <a href=\"http://egonw.posterous.com/nmrshiftdb-1006-contributions-and-counting\">1000th NMR spectrum</a> to the database, I thought it was time to\nfinally reinstall <a href=\"http://www.openlinksw.com/wiki/main/Main/VOSDownload\">Virtuoso</a>.</p>\n\n<p>There are precompiled binaries for <a href=\"https://launchpad.net/~wdaniels/+archive/ppa\">Ubuntu</a> and <a href=\"http://bugs.debian.org/cgi-bin/bugreport.cgi?bug=508048\">Debian</a>,\nbut Michel encouraged me to use version 6 when <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/06/26/michel-dumontier-at-uppsala-university.html\">he visited us <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.\nAnd so I compiled and install <a href=\"https://sourceforge.net/projects/virtuoso/files/virtuoso-devel/6.0.0-TP1/\">6.0.0.TP1</a> on the public server, while I do have the\nbinary debs for 5.0.12 on my laptop. With some basic Apache magic, I hooked up the SPARQL end point of the server to the web:</p>\n\n<div class=\"language-xml highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code><span class=\"nt\">&lt;Proxy</span> <span class=\"err\">/nmrshiftdb/sparql</span><span class=\"nt\">&gt;</span>\n  RewriteEngine On\n  Allow from all\n  ProxyPass        http://localhost:8890/sparql\n  ProxyPassReverse http://localhost:8890/sparql\n<span class=\"nt\">&lt;/Proxy&gt;</span>\n</code></pre></div></div>\n\n<p>Nice thing about this is, that I can set up multiple servers, allowing me to keep incompatibly licensed data sets apart (see\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/05/18/open-data-license-rights-aggregation.html\">Open Data: license, rights, aggregation, clean interfaces? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>), which is\nthe same approach Bio2RDF is taking.</p>\n\n<p>The <a href=\"http://pele.farmbio.uu.se/nmrshiftdb/sparql\">end point</a> now offers about <a href=\"http://pele.farmbio.uu.se/nmrshiftdb/sparql?default-graph-uri=&amp;query=SELECT+count%28*%29+WHERE+{\\%0D%0A++%3Fs+%3Fp+%3Fo+.%0D%0A}&amp;format=text%2Fhtml&amp;debug=on\">278887</a>\ntriples, but this will soon rise as I make more content from the database available in the original SQL database. The data is from the\n<a href=\"https://sourceforge.net/projects/nmrshiftdb/files/nmrshiftdb/1.3.3/\">1.3.3 release</a> by <a href=\"http://www.steinbeck-molecular.de/steinblog/\">Chris</a>’\nteam, and does not include my 1000th spectrum.</p>\n\n<p>Getting the data into the database was not trivial either. The documentation suggests WebDAV, and that indeed worked for me once, after\nusing the <a href=\"http://www.snee.com/bobdc.blog/2009/02/getting-started-using-virtuoso.html\">curl approach suggested here</a>. But upon a second upload, it\ndid again not enter the store. The ultimate solution was to use the iSQL interface, with the following SQL</p>\n\n<div class=\"language-plaintext highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code>DB.DBA.RDF_LOAD_RDFXML_MT(\n  file_to_string_output('/tmp/nmrshiftdb.rdf'), '',\n  'http://pele.farmbio.uu.se/nmrshiftdb'\n);\n</code></pre></div></div>\n\n<p>Scientifically, this progress is not overly interesting, although it makes it very clear that you really should not have to be happy with proprietary\nand non-semantic formats for anything. But, to me, this is mostly a technological success of great importance: I can now share really large sets of\nRDF data.</p>\n\n<p>Querying this data is a simple with SPARQL, and the results are available in various formats, such as JSON, which makes it easy to integrate in\nthird-party applications or <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/09/02/google-wave-robot-for-cdk-functionality.html\">Google Wave robots <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\n(did I hear someone say <a href=\"http://nmrshifty.appspot.com/\">NMRShifty</a>?). As I have <a href=\"http://chem-bla-ics.blogspot.com/search?q=sparql\">blogged before</a>,\nSPARQL is an excellent tool to aggregate scientific data prior to data analysis. And I will demo more interesting queries later this month.</p>",
      "summary": "About 6 months ago I reported about my efforts to RDF-ize the data from the NMRShiftDB. Since then, time was consumed by many other things, but now that Bioclipse can query SPARQL end points, that I want to contribute the triple set (it is GNU FDL-licensed) to Bio2RDF, that a student started working in my group (now larger than just me :) on reasoning on life sciences data, and that I recently contributed my 1000th NMR spectrum to the database, I thought it was time to finally reinstall Virtuoso.",
      
      "date_published": "2009-09-04T00:00:00+00:00",
      "date_modified": "2026-09-27T00:00:00+00:00",
      "tags": ["rdf","sparql","nmrshiftdb","cheminf"],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
