{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2024/02/13/wikidata-subsetting.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/57rv7-5m756",
      "url": "https://chem-bla-ics.linkedchemistry.info/2024/02/13/wikidata-subsetting.html",
      "title": "New paper: &quot;Wikidata subsetting: approaches, tools, and evaluation&quot;",
      "content_html": "<p>Just before the end of the year, the <em>Wikidata subsetting: approaches, tools, and evaluation</em> paper\nby Seyed Amir Hosseini Beghaeiraveri <em>et al.</em> got published (doi:<a href=\"https://doi.org/10.3233/SW-233491\">10.3233/SW-233491</a>).\nI am really excited our group (i.e.\n<a href=\"https://orcid.org/0000-0002-8399-8990\">Ammar</a> and <a href=\"https://orcid.org/0000-0001-8449-1318\">Denise</a>)\nhas been able to contribute to this. I think it also is a great example\nof the power of hackathons to bring together people.</p>\n\n<p>To me, subsetting of Wikidata (or any large knowledge graph) is important for a couple of reasons.\nFirst, there can be practical reasons. Scholia, for example, is computationally expensive, and the idea\nwe explore in the Alfred P. Sloan Foundation grant for Scholia (doi:<a href=\"https://doi.org/10.3897/rio.5.e35820\">10.3897/rio.5.e35820</a>)\nwas that a subset of Wikidata would make it more performant and potentially\nmore environmental-friendly.</p>\n\n<p>A second reason is more about the scientific process. When doing an analysis and when you want to make\nthe reasoning transparent, you want to share the analyzed data as part of the research output (basically, the “data”).\nFor example, the data may have undergone some curation, or you combined data from two or more different\nsources. And you will want to share this as part of the scientific process. Resharing a full dump\nof the larger knowledge base would not be practical for at least two reasons: duplication of huge data,\nand a lot of unrelated content makes it hard for peers to find the bits of interest to the study.</p>\n\n<p>Subsetting may be useful here. This paper evaluates a number of different subsetting approaches.\nMyself, I am particularly excited about the idea that we can take a shape expression (e.g. <a href=\"https://shex.io\">ShEx</a>)\nas input. I still love the idea that I take the SPARQL queries in my analyses, convert that into\nshapes automatically, and then get a subet that returns the exact same results as the query would\non the full dataset.</p>\n\n<h4>References</h4>\n<div class=\"csl-bib-body\">\n    <div class=\"csl-entry\">Hosseini Beghaeiraveri, S. A., Labra Gayo, J. E., Waagmeester, A., Ammar, A., Gonzalez, C., Slenter, D., Ul-Hasan, S., Willighagen, E., McNeill, F., &#38; Gray, A. J. G. (2023). Wikidata subsetting: Approaches, tools, and evaluation. <i>Semantic Web</i>, 1–27. https://doi.org/10.3233/sw-233491 <a href=\"https://doi.org/10.3233/SW-233491\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.3233/SW-233491\">Scholia</a></div>\n    <div class=\"csl-entry\">Rasberry, L., Willighagen, E., Nielsen, F., &#38; Mietchen, D. (2019). Robustifying Scholia: paving the way for knowledge discovery and research assessment through Wikidata. <i>Research Ideas and Outcomes</i>, <i>5</i>. https://doi.org/10.3897/rio.5.e35820 <a href=\"https://doi.org/10.3897/RIO.5.E35820\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.3897/RIO.5.E35820\">Scholia</a></div>\n  </div>",
      "summary": "Just before the end of the year, the Wikidata subsetting: approaches, tools, and evaluation paper by Seyed Amir Hosseini Beghaeiraveri et al. got published (doi:10.3233/SW-233491). I am really excited our group (i.e. Ammar and Denise) has been able to contribute to this. I think it also is a great example of the power of hackathons to bring together people.",
      "image": "https://chem-bla-ics.linkedchemistry.info/assets/images/wikidata_subsetting_features.png",
      "date_published": "2024-02-13T00:00:00+00:00",
      "date_modified": "2024-02-13T00:00:00+00:00",
      "tags": ["wikidata","scholia"],
      "_references": [
        
          
          
            { "url": "https://doi.org/10.3233/SW-233491", "doi": "10.3233/SW-233491"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.3897/RIO.5.E35820", "doi": "10.3897/RIO.5.E35820"
             }
            
          
        ],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
