{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2021/02/16/downloading-all-currently-released.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/erxg2-sm862",
      "url": "https://chem-bla-ics.linkedchemistry.info/2021/02/16/downloading-all-currently-released.html",
      "title": "Downloading all currently released BridgeDb identifier mapping databases",
      "content_html": "<p>The <a href=\"https://bridgedb.github.io/\">BridgeDb</a> project (doi:<a href=\"https://doi.org/10.1186/1471-2105-11-5\">10.1186/1471-2105-11-5</a>)\n(and <a href=\"https://elixir-europe.org/platforms/interoperability/rirs\">ELIXIR recommended interoperability resource</a>) has several\naims, all around identifier mapping:</p>\n\n<ul>\n  <li>provide a Java API for identifier mapping</li>\n  <li>provide ID mappings (two flavors: with and without semantic meaning)</li>\n  <li>provide services (<a href=\"https://www.bioconductor.org/packages/release/bioc/html/BridgeDbR.html\">R package</a>,\n<a href=\"http://webservice.bridgedb.org/\">OpenAPI webservice</a>)</li>\n  <li>track the history of identifiers</li>\n</ul>\n\n<p>The last one is more recent and two aspects are under development here: secondary identifiers and dead identifiers. More\nabout that in some future post. About the first and the third I am also not going to tell much in this post. Just follow the\nabove links.</p>\n\n<p>I do want to say something in this post about the actually identifier mapping databases, in particular those we distribute as\nApache Derby files, the storage format used by the Java libraries. These are the files you download if you want mapping databases\nfor <a href=\"https://pathvisio.github.io/\">PathVisio</a> (doi:<a href=\"https://doi.org/10.1371/journal.pcbi.1004085\">10.1371/journal.pcbi.1004085</a>).\nBridgeDb has mapping files for various things and some example databases the data it maps between:</p>\n\n<ol>\n  <li>genes and proteins: Ensembl, UniProt, NCBI Gene</li>\n  <li>metabolites; HMDB, ChEBI, LIPID MAPS, Wikidata, CAS</li>\n  <li>publications: DOI, PubMed</li>\n  <li>macromolecular complexes: Complex Portal, Wikidata</li>\n</ol>\n\n<p>The BridgeDb API is agnostic to the things it can map identifiers for.</p>\n\n<p><strong>Downloading mapping files</strong>:\nBridgeDb has an <a href=\"https://bioschemas.org/\">BioSchemas</a>-powered\n<a href=\"https://bridgedb.github.io/data/gene_database/\">web page with an overview of the latest released mapping files</a>.\nIt looks like this:</p>\n\n<p><img src=\"/assets/images/bridgedbDownloadsImage.png\" alt=\"\" /></p>\n\n<p>This webpage is the result from the cyber attack in late 2019, disrupting a good bit of the infrastructure. This is why we\nrenewed the website, including the download page. The new page actually is hosted <a href=\"https://github.com/bridgedb/data\">on GitHub as a Markdown file</a>,\nbut this is where things get interesting. The Markdown file is actually autogenerated from a JSON file with all the info. Everything,\nincluding the BioSchemas annotation is created from that. Basically, JSON gets converted into Markdown (with a custom script), which\ngets converted into HTML by a GitHub Action/Pages. So, when someone releases a new mapping file on Zenodo or Figshare, they only have\nto send me a pull request with updated JSON file.</p>\n\n<p>Now, previously, downloading all released mapping files, for example for the BridgeDb webservice, was a bit complicated. The\ninformation was a HTML file generated by the webserver for a folder. No metadata. Nuno wrote code to extract the relevant info\nand download all the files. However, since the information is now available in a public JSON file, it is a lot easier. The\nfollowing code uses wget and jq, two tools readily available on the popular operating systems. Have fun!</p>\n\n<div class=\"language-shell highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code><span class=\"o\">!</span>/bin/bash\n\nwget <span class=\"nt\">-nc</span> https://bridgedb.github.io/data/gene.json\nwget <span class=\"nt\">-nc</span> https://bridgedb.github.io/data/corona.json\nwget <span class=\"nt\">-nc</span> https://bridgedb.github.io/data/other.json\n\njq <span class=\"nt\">-r</span> <span class=\"s1\">'.mappingFiles | .[] | \"\\(.file)=\\(.downloadURL)\"'</span> gene.json <span class=\"o\">&gt;</span> files.txt\njq <span class=\"nt\">-r</span> <span class=\"s1\">'.mappingFiles | .[] | \"\\(.file)=\\(.downloadURL)\"'</span> corona.json <span class=\"o\">&gt;&gt;</span> files.txt\njq <span class=\"nt\">-r</span> <span class=\"s1\">'.mappingFiles | .[] | \"\\(.file)=\\(.downloadURL)\"'</span> other.json <span class=\"o\">&gt;&gt;</span> files.txt\n\n<span class=\"k\">for </span>FILE <span class=\"k\">in</span> <span class=\"si\">$(</span><span class=\"nb\">cat </span>files.txt<span class=\"si\">)</span>\n<span class=\"k\">do\n  </span>readarray <span class=\"nt\">-d</span> <span class=\"o\">=</span> <span class=\"nt\">-t</span> splitFILE<span class=\"o\">&lt;&lt;&lt;</span> <span class=\"s2\">\"</span><span class=\"nv\">$FILE</span><span class=\"s2\">\"</span>\n  <span class=\"nb\">echo</span> <span class=\"k\">${</span><span class=\"nv\">splitFILE</span><span class=\"p\">[0]</span><span class=\"k\">}</span>\n  wget <span class=\"nt\">-nc</span> <span class=\"nt\">-O</span> <span class=\"k\">${</span><span class=\"nv\">splitFILE</span><span class=\"p\">[0]</span><span class=\"k\">}</span> <span class=\"k\">${</span><span class=\"nv\">splitFILE</span><span class=\"p\">[1]</span><span class=\"k\">}</span>\n<span class=\"k\">done</span>\n</code></pre></div></div>\n\n<p>Actually, while writing this blog post, I notice the code can be further simplified.</p>\n\n<h4>References</h4>\n<div class=\"csl-bib-body\">\n    <div class=\"csl-entry\">Kutmon, M., van Iersel, M. P., Bohler, A., Kelder, T., Nunes, N., Pico, A. R., &#38; Evelo, C. T. (2015). PathVisio 3: An Extendable Pathway Analysis Toolbox. <i>PLOS Computational Biology</i>, <i>11</i>(2), e1004085. https://doi.org/10.1371/journal.pcbi.1004085 <a href=\"https://doi.org/10.1371/journal.pcbi.1004085\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1371/journal.pcbi.1004085\">Scholia</a></div>\n    <div class=\"csl-entry\">van Iersel, M. P., Pico, A. R., Kelder, T., Gao, J., Ho, I., Hanspers, K., Conklin, B. R., &#38; Evelo, C. T. (2010). The BridgeDb framework: standardized access to gene, protein and metabolite identifier mapping services. <i>BMC Bioinformatics</i>, <i>11</i>(1). https://doi.org/10.1186/1471-2105-11-5 <a href=\"https://doi.org/10.1186/1471-2105-11-5\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/1471-2105-11-5\">Scholia</a></div>\n  </div>",
      "summary": "The BridgeDb project (doi:10.1186/1471-2105-11-5) (and ELIXIR recommended interoperability resource) has several aims, all around identifier mapping:",
      "image": "https://chem-bla-ics.linkedchemistry.info/assets/images/bridgedbDownloadsImage.png",
      "date_published": "2021-02-16T00:00:00+00:00",
      "date_modified": "2025-02-16T00:00:00+00:00",
      "tags": ["bridgedb","json"],
      "_references": [
        
          
          
            { "url": "https://doi.org/10.1186/1471-2105-11-5", "doi": "10.1186/1471-2105-11-5"
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1371/journal.pcbi.1004085", "doi": "10.1371/journal.pcbi.1004085"
             }
            
          
        ],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
