{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2009/05/18/open-data-license-rights-aggregation.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/hvqxm-xnq47",
      "url": "https://chem-bla-ics.linkedchemistry.info/2009/05/18/open-data-license-rights-aggregation.html",
      "title": "Open Data: license, rights, aggregation, clean interfaces?",
      "content_html": "<p>A <a href=\"http://blog.openwetware.org/scienceintheopen/2009/05/15/a-breakthrough-on-data-licensing-for-public-science/\">recent post</a> by\n<a href=\"http://blog.openwetware.org/scienceintheopen/\">Cameron</a> on his visit last week with <a href=\"http://wwmm.ch.cam.ac.uk/blogs/adams/\">Nico</a>,\n<a href=\"http://wwmm.ch.cam.ac.uk/blogs/murrayrust/\">Peter</a> and <a href=\"http://wwmm.ch.cam.ac.uk/blogs/downing/\">Jim</a>, discussed\n<a href=\"http://en.wikipedia.org/wiki/Open_data\">Open Data</a> licensing. This lead to an interesting discussion on these matters, and\nquestions by me on why people care so much about only public domain data (or licensed with\n<a href=\"http://www.opendatacommons.org/licenses/pddl/1.0/\">PDDL</a> or <a href=\"http://wiki.creativecommons.org/CC0\">CC0</a>).</p>\n\n<p>Open licensing for data has not as much matured as for software, and international law seems to be more confusing about the\nissues. I guess that is because data aggregation has been around for way before the computer era. The PDDL and CC0 both try to\novercome this fuzziness. But there is another issue we need to keep in mind. A lot of useful Data was aggregated and made Open\n<em>before</em> these licenses came about, and use, for example, the <a href=\"http://www.gnu.org/copyleft/fdl.html\">GNU FDL</a> license, such as\nthe <a href=\"http://www.nmrshiftdb.org/\">NMRShiftDB</a>.</p>\n\n<h2 id=\"rights\">Rights</h2>\n\n<p>Right now, there are two Open Data camps, much like the BSD-vs-GPL wars in Open Source: one that believes in waiving any rights\non the Data, indicating that facts are free; others that believe that data must be protected to not be eaten by big companies\nand lost to the community (e.g. <a href=\"http://friendfeed.com/onssolubility/cf6afd52/should-we-contribute-solubility-data-to\">the WolframAlpha arragnements are suspect</a>).</p>\n\n<p>Of course, both camps are not that far apart, and both believe Open is important. Interestingly, there are some noteworthy\ndifferences with the Open Source wars. I see parallels between the two, which details an important difference: Open Source has\nalgorithms (uncopyrightable) and implementations (copyrightable); Open Data has Data (uncopyrightable) and aggregation\n(copyrightable). Open Source talks mostly about the implementation, not the algorithm; it’s Open Source, not Open Algorithms\nafter all. In cheminformatics it is even often the case that the algorithms are not even specified and that there only truly\nis source.</p>\n\n<p>However, Open Data in title does not make distinction. Data is fairly cheap and acquisition can be automated and computerized;\nAggregation, on the other hand, requires human involvement: curation and thinking about data models, etc. This is where added\nvalue is. Consider an assigned NMR spectrum or the raw data returned from the spectrometer.</p>\n\n<p>It is this added value that people want to protect, not the data itself. I think.</p>\n\n<h2 id=\"aggregation\">Aggregation</h2>\n\n<p>One important argument that tend to show up when people argument for PDDL and CC0 is that it makes data aggregation easier.\nThis is most certainly true: if you can do whatever you like with a blob of data, that also means aggregate with any other\nblob of data. However, copyleft licenses, like the GNU FDL, require the aggregation to have a compatible license too. It is\nthe license incompatibilities that make this impossible. Or … ?</p>\n\n<p>Open Source has matured to such a point that it is fairly clear what the intended behaviour is, regarding derivatives. An\naggregation of software (typically refered to as a distribution) is only a derivative under certain conditions. This makes\nit possible to run proprietary software on top of GNU/Linux, which uses the GNU GPL but does not require software to run on\ntop of it to be GPL too. Unless… unless, not a clear well-defined interface has been used, indicating a strong dependency.\nNow, surely, these things have not been confirmed to match actual law in court, but the intentions are clear.</p>\n\n<h2 id=\"clean-data-interfaces\">Clean Data Interfaces?</h2>\n\n<p>Now, if we would translate this to Open Data, would there be the equivalent of a clean interface? Can we build a data\ndistribution with data of various licenses? I think we can! I am not a lawyer and please consider this an invitation\nto discuss these matters…</p>\n\n<p>Let’s start simlpe… if I put a GNU FDL image in this blog, by linking to it with a open, free, clean HTML interface\n(<code class=\"language-plaintext highlighter-rouge\">&lt;img src=\"\"/&gt;</code>), would that make my blog GNU FDL too? I don’t think so. Surely, I would need to list copyright owner,\nand actually would be required to put the GNU FDL in my blog too, but hope linking to the license text would suffice too.\n(Let’s skip fair use at this moment, and assume the use goes beyond fair use). Question: am I not using a clean interface,\nand would this not make the image’s license no infect my blog?</p>\n\n<p>A more difficult example, consider <a href=\"http://rdf.openmolecules.net/\">rdf.openmolecules.net</a>, which surely aggregated facts,\nincluding data from the NMRShiftDB and <a href=\"http://dbpedia.org/\">DBPedia</a>. I am using a unique identifiers here, the NMRShiftDB\ncompound ID, and the DBPedia URL, which surely is GNU FDL, and use this to make a <code class=\"language-plaintext highlighter-rouge\">&lt;owl:sameAs&gt;</code> statement. Again, please do\nnot consider fair use, which this certainly is. But, let’s say I put in some more DBPedia and NMRShiftDB data in this\naggregation. The GNU FDL data on rdf.openmolecules.net would be separate RDF blocks, with proper dc:license, dc:author\nannotation. But the block would be part of a larger aggregation. The clean interface here is\n<a href=\"http://en.wikipedia.org/wiki/Resource_description_framework\">Resource Description Framework</a>.</p>\n\n<p>This second case does not only affect my rdf.openmolecules.net website, but, for example, <a href=\"http://bio2rdf.org/\">bio2rdf.org</a>\nis also in the same situation and aggregated and distribute DBPedia’s GNU FDL data (e.g.\n<a href=\"http://bio2rdf.org/searchns/dbpedia/hexokinase\">hexinanose</a>. Does that make the\nwhole of bio2rdf database GNU FDL. They too use RDF as clean interface.</p>\n\n<h2 id=\"call-for-discussion\">Call for Discussion</h2>\n\n<p>Despite what one of the two camps like to see, the mere fact of added value when making data aggregations will keep\ncopyleft license stay around, and instead of trying to convince everyone of the virtues of PDDL- and CC0-like licenses,\nwe should think about to what extend it really matters.</p>\n\n<p>I can do my data analysis with data sources of various licenses. I can search and retrieve data from various sources\nwith various licenses. What obstacles are really there that disallow us to do science? Do the data interfaces we have\nnow not provide enough technical means to address the license incompatibilities? They have in Open Source, why would\nthat not apply to Open Data too?</p>",
      "summary": "A recent post by Cameron on his visit last week with Nico, Peter and Jim, discussed Open Data licensing. This lead to an interesting discussion on these matters, and questions by me on why people care so much about only public domain data (or licensed with PDDL or CC0).",
      
      "date_published": "2009-05-18T00:00:00+00:00",
      "date_modified": "2009-05-18T00:00:00+00:00",
      "tags": ["opendata","nmrshiftdb","rdf","dbpedia","bio2rdf"],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
