A bit over half a year ago, Markus Englund wrote about copy/paste errors in scientific datasets. I believe that many of them are just because spreadsheets make these things easy, but when you look at the examples, and realize that routinely datasets are found with particularly non-random distributions if data, we know that some of this is fraud. Markus’ post is a must read.

It me also made me realize something else. We have by now a good practise of tracking retracted articles, and platforms to discuss articles openly, like PubPeer. The second can be used for datasets too, when they are deposited in DOI-providing resources, like Zenodo or Figshare.

Fraudulent data can certainly result in papers that describe the data to be retracted. But retracting the publication does not mean the paper does not get cited anymore, see doi:10.1007/s11192-020-03631-1 or Scholia for retraction data. But data that starts having a live on its own after the deposit in some archive may worsen this problem. After all, our practises for sharing retraction flags and concerns are even without this substandard.

Should we not also start taking more care about the data we publish? Track concerns about them? Maybe even a way to retract datasets, when fraud was in play?