{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2025/08/06/archiving-but-not-really.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/vwd81-p8z85",
      "url": "https://chem-bla-ics.linkedchemistry.info/2025/08/06/archiving-but-not-really.html",
      "title": "Archiving, but not really",
      "content_html": "<p><a href=\"https://sauropods.win/@mike\">Mike Taylor</a> wrote up <a href=\"https://doi.org/10.59350/svpow.24000\">a post</a> about the various things a journal article is doing,\nthe first being <em>a scientific report</em>. We put a lot of money in establishing a scientific track record. In the past 30 years\nhow we publish our research and how we archive it has changed significantly. If you read my blog more often, you know I have\nbeen critical of the performance of many publishers. Springer Nature was so disappointing that after 5 years I\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2021/06/11/conflict-of-interest-or-why-i-am.html\">stepped down</a>\nas Editor-in-Chief (of two) of the <a href=\"https://en.wikipedia.org/wiki/Journal_of_Cheminformatics\">Journal of Cheminformatics</a>.\nThere is so much that must be <a href=\"https://chem-bla-ics.linkedchemistry.info/2024/09/16/publishing.html\">done better</a>.</p>\n\n<p>But in the most recent iteration, triggered by some work for <a href=\"https://www.wikipathways.org/\">WikiPathways</a>, I was using\n<a href=\"https://europepmc.org/\">Europe PMC</a> to find articles that\nmention <em>WikiPathways</em> and then search in the full text for the string <code class=\"language-plaintext highlighter-rouge\">WP</code>, as a trigger for the possible mention of\nWikiPathways pathway identifiers, which look like <code class=\"language-plaintext highlighter-rouge\">WP4846</code>. The use of <em>compact (resource) identifiers</em>\n(see doi:<a href=\"https://doi.org/10.1038/sdata.2018.29\">10.1038/sdata.2018.29</a>) is minimal, but at least some articles use identifiers.</p>\n\n<p>That allows me to extend our WikiPathways knowledge graph of <a href=\"https://www.wikipathways.org/browse/citedin\">articles citing specific pathways</a>.\nAt the time of writing, we collected 2509 citations from 440 different articles to 883 different pathways. Now,\nI want to blog about that more, but it’s related to an observation.</p>\n\n<h2 id=\"information-loss\">Information loss</h2>\n<p>Now, back in the late ninities I learned about GNU/Linux and after playing with Red Hat and Suse, I settled for Debian.\nOne of the things I learned is that, generally, information corruption (like data loss) is an absolute red flag, a no-go,\na total showstopper.</p>\n\n<p>And then we have this in publishing, the one area where data corruption must also be a no-go:</p>\n\n<p><img src=\"/assets/images/imageResolutionLoss.png\" alt=\"\" /></p>\n\n<p>In this image, the left side shows a screenshot of the publisher version of the article and on the right side\nthe version in <a href=\"https://pmc.ncbi.nlm.nih.gov/\">Pubmed Central</a> (PMC). PMC has been an important project to archive full text versions of articles:</p>\n\n<blockquote>\n  <p>11.2 million articles are archived in PMC.</p>\n</blockquote>\n\n<p>So, this is <strong>really bad</strong>! The archived version is not really useful. As a human I already struggle to read the\ndegraded image, let alone an algorithm.</p>\n\n<p>Does that matter? Yes, projects like the awesome\n<a href=\"https://pfocr.wikipathways.org/\">Pathway Figure OCR</a> (see doi:<a href=\"https://doi.org/10.1186/s13059-020-02181-2\">10.1186/s13059-020-02181-2</a>)\ndepend on images to be FAIR enough to extract information. (Side note: yes, these images should be vector\ngraphics, but commercial publishers decided about twenty years ago that they could not care enough.)</p>\n\n<p>At this moment, I do not know where the information is lost. Maybe PubMed Central is storing the images in a low\nresolution. Maybe the publisher provides PMC with a low resolution image. But to me, this must be solved as soon\nas possible. This is utterly unacceptable.</p>\n\n<p>I wonder what the authors of the article (doi:<a href=\"https://doi.org/10.1186/s13287-025-04166-z\">10.1186/s13287-025-04166-z</a>)\nI took as example think of this.</p>\n\n<h4>References</h4>\n<div class=\"csl-bib-body\">\n    <div class=\"csl-entry\">Hanspers, K., Riutta, A., Summer-Kutmon, M., &#38; Pico, A. R. (2020). Pathway information extracted from 25 years of pathway figures. <i>Genome Biology</i>, <i>21</i>(1). https://doi.org/10.1186/s13059-020-02181-2 <b>[cito:citesAsRecommendedReading]</b> <a href=\"https://doi.org/10.1186/s13059-020-02181-2\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/s13059-020-02181-2\">Scholia</a></div>\n    <div class=\"csl-entry\">Taylor, M. (2025). Journal articles are trying to do six things at once — no wonder they’re unreadable. In <i>Front Matter</i>. Front Matter. https://doi.org/10.59350/svpow.24000 <b>[cito:citesAsRecommendedReading]</b> <a href=\"https://doi.org/10.59350/svpow.24000\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.59350/svpow.24000\">Scholia</a></div>\n    <div class=\"csl-entry\">Wimalaratne, S. M., Juty, N., Kunze, J., Janée, G., McMurry, J. A., Beard, N., Jimenez, R., Grethe, J. S., Hermjakob, H., Martone, M. E., &#38; Clark, T. (2018). Uniform resolution of compact identifiers for biomedical data. <i>Scientific Data</i>, <i>5</i>(1). https://doi.org/10.1038/sdata.2018.29 <b>[cito:obtainsBackgroundFrom]</b> <a href=\"https://doi.org/10.1038/sdata.2018.29\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1038/sdata.2018.29\">Scholia</a></div>\n    <div class=\"csl-entry\">Yuan, F., Liu, J., Zhong, L., Liu, P., Li, T., Yang, K., Gao, W., Zhang, G., Sun, J., &#38; Zou, X. (2025). Enhanced therapeutic effects of hypoxia-preconditioned mesenchymal stromal cell-derived extracellular vesicles in renal ischemic injury. <i>Stem Cell Research &#38;amp; Therapy</i>, <i>16</i>(1). https://doi.org/10.1186/s13287-025-04166-z <b>[cito:describes]</b> <a href=\"https://doi.org/10.1186/s13287-025-04166-z\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.1186/s13287-025-04166-z\">Scholia</a></div>\n  </div>",
      "summary": "Mike Taylor wrote up a post about the various things a journal article is doing, the first being a scientific report. We put a lot of money in establishing a scientific track record. In the past 30 years how we publish our research and how we archive it has changed significantly. If you read my blog more often, you know I have been critical of the performance of many publishers. Springer Nature was so disappointing that after 5 years I stepped down as Editor-in-Chief (of two) of the Journal of Cheminformatics. There is so much that must be done better.",
      "image": "https://chem-bla-ics.linkedchemistry.info/assets/images/imageResolutionLoss.png",
      "date_published": "2025-08-06T00:00:00+00:00",
      "date_modified": "2025-08-06T00:00:00+00:00",
      "tags": ["publishing","europepmc"],
      "_references": [
        
          
          
            { "url": "https://doi.org/10.59350/svpow.24000", "doi": "10.59350/svpow.24000"
            , "cito":
              
              
                [ 
                  "citesAsRecommendedReading"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1186/s13059-020-02181-2", "doi": "10.1186/s13059-020-02181-2"
            , "cito":
              
              
                [ 
                  "citesAsRecommendedReading"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1186/s13287-025-04166-z", "doi": "10.1186/s13287-025-04166-z"
            , "cito":
              
              
                [ 
                  "describes"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.1038/sdata.2018.29", "doi": "10.1038/sdata.2018.29"
            , "cito":
              
              
                [ 
                  "obtainsBackgroundFrom"
                  
                 ]
              
             }
            
          
        ],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
