<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="https://chem-bla-ics.linkedchemistry.info/feed/by_tag/scholia.xml" rel="self" type="application/atom+xml" /><link href="https://chem-bla-ics.linkedchemistry.info/" rel="alternate" type="text/html" /><updated>2026-07-18T13:36:15+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/feed/by_tag/scholia.xml</id><title type="html">chem-bla-ics</title><subtitle>Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.</subtitle><author><name>Egon Willighagen</name></author><entry><title type="html">Rescuing Scholia #3: We did it!</title><link href="https://chem-bla-ics.linkedchemistry.info/2026/02/28/rescuing-scholia-3-we-did-it.html" rel="alternate" type="text/html" title="Rescuing Scholia #3: We did it!" /><published>2026-02-28T00:00:00+00:00</published><updated>2026-02-28T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2026/02/28/rescuing-scholia-3-we-did-it</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/02/28/rescuing-scholia-3-we-did-it.html"><![CDATA[<p>It was not a set up, when I openly <a href="https://chem-bla-ics.linkedchemistry.info/2025/12/08/rescuing-scholia.html">wondered if we would be able to rescue Scholia in time</a>.
I honestly did not know. Three weeks and some serious hacking by an international team later <a href="https://chem-bla-ics.linkedchemistry.info/2025/12/31/rescuing-scholia-2-getting-close.html">I was more optimistic</a>.
Actually, just before christmas, we started writing a <a href="https://www.swat4ls.org/">SWAT4HCLS 2026</a> demonstration abstract. This was accepted and
you can read the <em>Scholia 2026: Compliance with SPARQL 1.1</em> preprint <a href="https://github.com/WolfgangFahl/ScholiaGraphSplitPaper">here</a> and
<a href="https://commons.wikimedia.org/wiki/File:Scholia_2026_Compliance_with_SPARQL_1.1.pdf">here</a>.
This paper describes the work that had to be done, and I am deeply grateful to everyone who contributed with smaller or
bigger contributions (Daniel, Peter, Konrad, Johannes, Lars, Wolfgang, Hannah).
I am merely first author for the demo, and just another contributor to the long series of patches, in a
<a href="https://github.com/WDscholia/scholia/pull/2715">branch started by Prof. Hannah Bast</a>.</p>

<p>The work actually started long before that, with the <em>Robustifying Scholia</em> grant (see doi:<a href="https://doi.org/10.3897/rio.5.e35820">10.3897/rio.5.e35820</a>),
where we explored alternatives. The Wikidata graph (RDF) split has been long coming, and I can recommend
<a href="https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-02-17/Technology_report">this recent The Signpost article</a>
by <a href="https://disobey.net/@Bluerasberry">Lane</a> for a good overview. So, this would not have been possible with
<a href="https://github.com/WDscholia/scholia/graphs/contributors">the many people who contributed over the years</a>.
But this last sprint really made a difference.</p>

<p>The developments of the QLever software in the past year are very important, and the SPARQL endpoint we run now is live updated,
just like we knew from the Wikidata Query Service (WDQS). Recent improvement allowed us to replace all the Wikidata and Blazegraph
specific aspects of the SPARQL queries, and good discussions let to pragmatic approaches to keep localization features
Scholia had for displaying query results from Wikidata.</p>

<p>The work is not completed, however. All queries are SPARQL 1.1 now, but some can still be further optimized, and some still
need some fixing. For example, I still spot some QIDs here and there, instead of the localized labels that should be shown instead.
Also, we are actively looking in getting everything running again on WMF servers (see <a href="https://github.com/WDscholia/scholia/issues/2766">this overview issue</a>),
so that <em>scholia.toolforge.org</em> works again.</p>

<p>For now, however, please use <a href="https://qlever.scholia.wiki/">qlever.scholia.wiki</a>.</p>]]></content><author><name>Egon Willighagen</name></author><category term="scholia" /><category term="sparql" /><category term="swat4ls" /><category term="doi:10.3897/RIO.5.E35820" /><summary type="html"><![CDATA[It was not a set up, when I openly wondered if we would be able to rescue Scholia in time. I honestly did not know. Three weeks and some serious hacking by an international team later I was more optimistic. Actually, just before christmas, we started writing a SWAT4HCLS 2026 demonstration abstract. This was accepted and you can read the Scholia 2026: Compliance with SPARQL 1.1 preprint here and here. This paper describes the work that had to be done, and I am deeply grateful to everyone who contributed with smaller or bigger contributions (Daniel, Peter, Konrad, Johannes, Lars, Wolfgang, Hannah). I am merely first author for the demo, and just another contributor to the long series of patches, in a branch started by Prof. Hannah Bast.]]></summary></entry><entry><title type="html">Rescuing Scholia #2: getting closer</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/12/31/rescuing-scholia-2-getting-close.html" rel="alternate" type="text/html" title="Rescuing Scholia #2: getting closer" /><published>2025-12-31T00:00:00+00:00</published><updated>2025-12-31T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/12/31/rescuing-scholia-2-getting-close</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/12/31/rescuing-scholia-2-getting-close.html"><![CDATA[<p>Three weeks ago, I wrote a the post <a href="https://chem-bla-ics.linkedchemistry.info/2025/12/08/rescuing-scholia.html">Rescuing Scholia: will we make it in time?</a>,
where I sketched a future without <a href="https://scholia.toolforge.org/">Scholia</a>. Scholia, started
<a href="https://chem-bla-ics.linkedchemistry.info/2023/01/27/scholia-timeline.html">almost 10 years ago</a>
and I think it is worth keeping around longer.</p>

<p>Fortunately, it looks like we will have a working replacement in time before the
<a href="https://www.mediawiki.org/wiki/Wikidata_Query_Service">WDQS</a> instance with all the
<a href="https://wikidata.org/">Wikidata</a> triples in a single SPARQL endpoint goes down,
likely in a week or so (even tho we may be behind <a href="https://openalex.org/works?page=1&amp;filter=cites:w2767995756">the citation peak</a>).</p>

<p>The work of the past year helped, for exampe, making it easier to configure Scholia for a different
endpoint and the asynchronous loading of panels (reducing the stress on the SPARQL end point).
Already in September, Prof. <a href="https://github.com/hannahbast">Hannah Bast</a> started
<a href="https://github.com/WDscholia/scholia/pull/2715">a branch</a> for the transition and various
hackathons this autumn, and the work by <a href="https://github.com/KonradLinden">Konrad Linded</a>
who explored and addressed some of the hurdles to take. The tips and suggestions from
Hannah and <a href="https://github.com/RobinTF">RobinTF</a> really made a difference. And also a huge thanks
to <a href="https://orcid.org/0000-0001-9488-1870">Daniel</a> who kept relentlessly pushing this forward.</p>

<p>When I posted my <a href="https://chem-bla-ics.linkedchemistry.info/2025/12/08/rescuing-scholia.html">will we make it</a> post,
there was a demo instance and a spreadsheet showing the state of each query. The instance
showed no human-readable labels. This was because the WDQS <code class="language-plaintext highlighter-rouge">wikibase:label</code> service 
was used a lot, and there is no replacement for that. Getting labels for all relevant
items is possible, but makes the queries a lot heavier and made even more queries
run out of memory. Various solutions were <a href="https://github.com/ad-freiburg/scholia/issues/17">discussed</a>,
Finn indicated he <a href="https://github.com/ad-freiburg/scholia/issues/17#issuecomment-3605952951">preferred a macro solution</a>,
which <a href="https://github.com/ad-freiburg/scholia/pull/20/changes">Lars implemented</a>, and
saw some tweaks after that. Then followed a long series of patches by particularly
<a href="https://github.com/pfps">Peter</a> to update all the SPARQL queries to have them use
the new labels macro. But plenty of other things were fixed or newly implemented,
such as <a href="https://github.com/WolfgangFahl">Wolfgang</a>’s <a href="https://qlever.scholia.wiki/backend">/backend</a>
page.</p>

<p>So, with one week to go, we need your help: as the weekly
<a href="https://www.wikidata.org/wiki/Wikidata:Status_updates/2025_12_29">Wikidata Status Update</a>
already indicated:</p>

<blockquote>
  <p>this month’s Scholia hackathon has moved Scholia closer to its planned switch to a
QLever backend. Beta testers can assist by exploring the
<a href="https://qlever.scholia.wiki/">interim QLever-backed Scholia instance</a>
and <a href="https://github.com/WDscholia/scholia/issues">reporting any issues</a>.</p>
</blockquote>

<p>And thanks to <a href="https://github.com/Adafede">Adriano</a> and others who already have!</p>

<p>Now, we are not done yet. The real instance at <a href="https://scholia.toolforge.org/">scholia.toolforge.org</a>
has seen ridiculous abuse by scrapers (and the main instance is regularly unusable, to be honest),
and we have no idea the new setup is powerful enough. And we need to point to the new servers anyway.
So, plenty of work is left to be done in the next few days.</p>

<p>But we are getting close. So, please give <a href="https://qlever.scholia.wiki/">qlever.scholia.wiki</a>
a go, and let us know your observations. As <a href="https://en.wikipedia.org/wiki/Linus%27s_law">Linus’s law</a> writes:</p>

<blockquote>
  <p>Given enough eyeballs, all bugs are shallow.</p>
</blockquote>]]></content><author><name>Egon Willighagen</name></author><category term="wikidata" /><category term="scholia" /><category term="sparql" /><category term="rdf" /><summary type="html"><![CDATA[Three weeks ago, I wrote a the post Rescuing Scholia: will we make it in time?, where I sketched a future without Scholia. Scholia, started almost 10 years ago and I think it is worth keeping around longer.]]></summary></entry><entry><title type="html">Rescuing Scholia: will we make it in time?</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/12/08/rescuing-scholia.html" rel="alternate" type="text/html" title="Rescuing Scholia: will we make it in time?" /><published>2025-12-08T00:00:00+00:00</published><updated>2025-12-08T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/12/08/rescuing-scholia</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/12/08/rescuing-scholia.html"><![CDATA[<p>What <a href="https://chem-bla-ics.linkedchemistry.info/2023/01/27/scholia-timeline.html">started out in 2016 on Twitter</a> became a
<a href="https://meta.wikimedia.org/wiki/Coolest_Tool_Award/Full_history">(small) award winning</a>
<a href="https://chem-bla-ics.linkedchemistry.info/tag/scholia">decade long collaborative project</a>.
Unfortunately, the future is not clear. We are at odds if it will survice the growth of Wikidata
and in particularly the <a href="https://www.wikidata.org/wiki/Wikidata:SPARQL_query_service/WDQS_graph_split">SPARQL graph split</a>.
To be clear, the choice for Blazegraph initially worked great, but after it was bought by a big
company, developed halted. Very unfortunate for Wikidata. Unlike earlier, we no longer have funding, and rewriting Scholia
at this scale takes a good bit of effort. We already
<a href="https://chem-bla-ics.linkedchemistry.info/2025/04/20/the-april-2025-scholia-hackathon.html">held a few hackathons</a>.</p>

<p>So far, we have been able to continue to use a <em>legacy</em> SPARQL endpoint with all the data, but in exactly one month
that endpoint will be sunset. And we are <strong>not</strong> ready.</p>

<h2 id="rescuing-scholia">Rescuing Scholia</h2>
<p>Daniel and Lane have been leading an effort to rescue Scholia. The hackathons were part of this effort. It seems
that <a href="https://en.wikipedia.org/wiki/QLever">QLever</a> is the only route left. Earlier efforts to rewrite the more
than 350 Scholia SPARQL queries to support the graph split have basically failed. The complexity is far too high.
QLever, however, provides the full graph and since recently full SPARQL 1.1 support. That is also not enough to
reproduce the full Scholia functionality, but it seems to get us far.
Importantly, the data may not update as frequently as the <a href="https://www.mediawiki.org/wiki/Wikidata_Query_Service">WDQS</a>,
and that is another complexity to take into account. Particularly, all the 404 pages.</p>

<p>So, in the next weeks, we have to complete rewriting all those queries as queries that QLever can handle. A team
of people have done great work already, <a href="https://github.com/ad-freiburg/scholia/issues?q=is%3Aissue%20author%3AKonradLinden">including Konrad</a>.</p>

<p>I hope we make it in time.</p>]]></content><author><name>Egon Willighagen</name></author><category term="scholia" /><category term="wikidata" /><category term="rdf" /><category term="sparql" /><summary type="html"><![CDATA[What started out in 2016 on Twitter became a (small) award winning decade long collaborative project. Unfortunately, the future is not clear. We are at odds if it will survice the growth of Wikidata and in particularly the SPARQL graph split. To be clear, the choice for Blazegraph initially worked great, but after it was bought by a big company, developed halted. Very unfortunate for Wikidata. Unlike earlier, we no longer have funding, and rewriting Scholia at this scale takes a good bit of effort. We already held a few hackathons.]]></summary></entry><entry><title type="html">The Internet Journal of Chemistry</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/08/11/the-internet-journal-of-chemistry.html" rel="alternate" type="text/html" title="The Internet Journal of Chemistry" /><published>2025-08-11T00:00:00+00:00</published><updated>2025-08-11T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/08/11/the-internet-journal-of-chemistry</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/08/11/the-internet-journal-of-chemistry.html"><![CDATA[<p>The <a href="https://scholia.toolforge.org/topic/Q27211732">Internet Journal of Chemistry</a> (IJC, issn:1099-8292) was one of the first scientific journals to get
published on the world wide web (part of <em>the Internet</em>), see doi:<a href="https://doi.org/10.1080/00987913.2000.10764578">10.1080/00987913.2000.10764578</a>.
Issues were published from 1998 to 2004. But because it predates
systematic archiving of webpages by libraries, a lot is lost. The nature of the journal, however, makes it unique, and quite
a number of articles are cited a lot, and should be part of the <em>scientific record</em>.
But I soon realized it actually is quite hard to track down content of the journal. I knew some articles have been
<em>author accepted manuscripts</em> online. One of that was my own first (and single) author-article, self-archived on
Zenodo (doi:<a href="https://doi.org/10.5281/zenodo.1495470">10.5281/zenodo.1495470</a>), green open access style.</p>

<p>I wanted to see what I could recover, and here I describe what I did and what could be done next.</p>

<h2 id="a-list-of-all-articles">A list of all articles</h2>

<p>The first step is actually to create a list of all articles published in the IJC and collect as much metadata about
them as possible. With just over 100 articles, I decided to use Wikidata, as a machine-readable database, supporting the curation and reporting. I wanted at least
two independent sources, and for Wikidata, use public resources. That means, while Web of Science does have a list of
all articles, I only used this for validation, and <strong>not</strong> as information source. Instead, I used citations to IJC
articles and, of course, the Internet Archive (IA). It turns out <a href="https://web.archive.org/web/*/http://www.ijc.com/abstracts/*">a query like this</a>
does wonders (well, for the abstracts; I did not find full-texts archived on IA):</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>https://web.archive.org/web/*/http://www.ijc.com/abstracts/*
</code></pre></div></div>

<p>I found that all but one article had the abstract archived in the IA. Here’s <a href="https://web.archive.org/web/20000925050415/http://www.ijc.com/abstracts/abstract2n8.html">an example</a>:</p>

<p><img src="/assets/images/ia_ijc_abstract.png" alt="" /></p>

<p>This gave my a lot of information to add to Wikidata. Title, publication date, volume, article number, keywords, an absstract,
and, of course, the list of authors. Some authors I know personally, many I did not. But it did allow me to enter all
articles to Wikidata along with the authors and “author” (<a href="https://www.wikidata.org/wiki/Property:P50">P50</a>) or
“author name string” (<a href="https://www.wikidata.org/wiki/Property:P2093">P2093</a>).</p>

<h2 id="the-article-authors">The article authors</h2>

<p>It also turned out that multiple authors listed their IJC article on their public ORCID profile.
That greatly helped identification. I managed to <a href="https://w.wiki/Ezda">link many authors</a> to mostly existing Wikidata items:</p>

<p><img src="/assets/images/ijc_authors.png" alt="" /></p>

<p>I already mentioned that I used Wikidata to collect this information. Besides the <a href="https://scholia.toolforge.org/venue/Q27211732">interactive visualization with Scholia</a>,
it also gave me the option to track my progress with SPARQL queries. For example, <a href="https://w.wiki/Ezdf">this query</a> helped
me do that author FAIR-ification:</p>

<p><img src="/assets/images/ijc_sparql1.png" alt="" /></p>

<p>You can see here two columns with author information, one for P50 and the other for P2093. There is quite some
identification left to be done, and additional information is welcome:</p>

<p><img src="/assets/images/ijc_sparql2.png" alt="" /></p>

<h2 id="sources">Sources</h2>

<p>So, that brings us to this list of sources:</p>

<ul>
  <li>Internet Archive: abstracts and metadata</li>
  <li>ORCID profiles: ORCIDs of (some) authors</li>
  <li>Google Scholar: metadata and citations</li>
  <li>Web of Science: independent list for external validation</li>
</ul>

<p>Because there is plenty of work left to be done and I hope the collected information will further spread
in library collections, I added sources as much as possible. <a href="https://w.wiki/Em9i">This query</a> lists for all
articles the Web of Science identifier (recorded so that everyone can check the consistency), the link
to the Internet Archive-d abstract page, and a link to a known full text (five).</p>

<p>If you wonder, neither <a href="https://openalex.org/works?page=1&amp;filter=primary_location.source.id:s32147083">OpenAlex</a>
or <a href="https://europepmc.org/search?query=JOURNAL%3A%28%22Internet%20Journal%20of%20Chemistry%22%29">Europe PMC</a> have a full list.</p>

<h2 id="whats-next">What’s next?</h2>

<p>I do not have a formal training in archiving, but I am happy with the minimal viable metadata collection.
I know more can be done (and love to hear your pointers and suggestions): more author identies,
better coverage of keyword annotation, etc. But I think an important addition is adding citations
to and from the IJC articles are important. The journal predates efforts like the <a href="https://i4oc.org/">I4OC</a> and
<a href="https://opencitations.net/">Open Citations</a>, so I may have to manually recover citations from Google Scholar.
I will have to report on that later. But you can enjoy the citations that are
<a href="https://scholia.toolforge.org/venue/Q27211732#Citations">already there</a>. And now that we have sufficient metadata,
I can use this to find more full texts.</p>

<p>Btw, I have made contact with Prof. <a href="https://scholia.toolforge.org/author/Q28420106">Steven Bachrach</a>,
who founded the journal and was the Editor-in-Chief.</p>]]></content><author><name>Egon Willighagen</name></author><category term="publishing" /><category term="wikidata" /><category term="scholia" /><category term="doi:10.5281/ZENODO.1495470" /><category term="cito:citesAsEvidence:10.1080/00987913.2000.10764578" /><category term="europepmc" /><summary type="html"><![CDATA[The Internet Journal of Chemistry (IJC, issn:1099-8292) was one of the first scientific journals to get published on the world wide web (part of the Internet), see doi:10.1080/00987913.2000.10764578. Issues were published from 1998 to 2004. But because it predates systematic archiving of webpages by libraries, a lot is lost. The nature of the journal, however, makes it unique, and quite a number of articles are cited a lot, and should be part of the scientific record. But I soon realized it actually is quite hard to track down content of the journal. I knew some articles have been author accepted manuscripts online. One of that was my own first (and single) author-article, self-archived on Zenodo (doi:10.5281/zenodo.1495470), green open access style.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/ia_ijc.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/ia_ijc.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">PFAS in the blood of the Dutch population</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/07/06/pfas-in-the-blood-of-the-dutch-population.html" rel="alternate" type="text/html" title="PFAS in the blood of the Dutch population" /><published>2025-07-06T00:00:00+00:00</published><updated>2025-07-06T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/07/06/pfas-in-the-blood-of-the-dutch-population</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/07/06/pfas-in-the-blood-of-the-dutch-population.html"><![CDATA[<p>A recent report by the Dutch <a href="https://www.rivm.nl/">RIVM</a>, <em>PFAS in the blood of the Dutch population</em>
(doi:<a href="https://www.rivm.nl/bibliotheek/rapporten/2025-0094.pdf">10.21945/RIVM-2025-0094</a>), writes
that seven <a href="https://scholia.toolforge.org/chemical-class/Q648037">PFAS</a> compounds are found in blood samples
of all tested people. Another nine compounds are found in at least 1-in-10 people.
Because there is relevant data in the report on the 28 studied PFAS compound, I wanted to
have the report more FAIR than it is on the website. Why this report? Well, the chemistry and the
history is fascinating and brutal (I like <a href="https://www.youtube.com/watch?v=SC2eSujzrUY">this Veritasium video</a>).</p>

<p>The history tells me that our society may sound woke and leftish, in reality it is a continous fight
for basic human rights. (Something that plenty have been saying for years.)
In this case, a healtht life is the human right.</p>

<p>So, what can I do to make this report more FAIR?</p>

<h2 id="findable-in-wikidata">Findable in Wikidata</h2>

<p>Since this report has been <a href="https://news.google.com/search?q=PFAS%20in%20the%20blood%20of%20the%20Dutch%20population&amp;hl=en-US&amp;gl=US&amp;ceid=US%3Aen">mentioned in the news</a>,
it clearly is notable. The simplest thing to do is thus to just add it <a href="https://www.wikidata.org/wiki/Wikidata:Main_Page">Wikidata</a>.
Because the DOI of the report had not been recorded yet, I could not let <a href="https://scholia.toolforge.org/">Scholia</a>
do it for me. But doing it manually is only a bit more work: <a href="https://www.wikidata.org/wiki/Q135222054">Q135222054</a>.
The provided metadata <a href="https://www.rivm.nl/en/news/first-nationwide-study-into-pfas-in-blood">on the RIVM website</a>
is minimal.</p>

<p>But we can do more. Particularly, because I want people to find this report when they look info knowledge
about the 28 studied chemicals, I added <a href="https://www.wikidata.org/wiki/Q135222054#P921">main subject</a> annotation
using the information in <em>Table 1</em> in the report. Using Scholia and the CAS registry number in the table,
I crosscheck the information in Wikidata is consistent with the report (and visa versa). It was.
I then added the Dutch name and acronym for most of them. Some already had the name as in the Table.
That gives us a nice “Topic scores” plot for <a href="https://scholia.toolforge.org/work/Q135222054">the Scholia page of the report</a>:</p>

<p><img src="/assets/images/pfas_report.png" alt="" /></p>

<p>The central PFAS bubble is also only one <em>main subject</em> but larger because many the specific PFAS compounds
are subclassing PFAS. And you may also note many smaller bubbles. These actually come from <em>main subject</em>
annotations of articles cited from the report. Because I added a few of them too. Not all, because many are
not in Wikidata (yet).</p>

<h2 id="findable-in-wikipathways">Findable in WikiPathways</h2>

<p>But since 16 of these compounds are readily found in human blood samples, that is handy knowledge when
doing metabolomics (on blood samples). Or (and I leave that to later blog post), we can map the experimental
data for Dordrecht versus the rest of The Netherlands to the PFAS compounds. That is relevant to research
by <a href="https://vhp4safety.nl">VHP4Safety</a>. There are many ways to see if you have PFAS in your dataset,
but since we have many controlled lists of genes in metabolites, I added one for common PFAS in human
blood samples. Well, the 16 common in Dutch blood samples:</p>

<p><img src="/assets/images/pfas_wikipathways.png" alt="" /></p>

<p>Each <em>metabolite</em> here is annotated with their Wikidata identifier, allowing us to map experimental
data on top of it. And we get links out to other databases almost for free:</p>

<p><img src="/assets/images/pfas_wikipathways_outlinks.png" alt="" /></p>

<p>And the link to Wikidata actually links to Scholia, so for the PFOA in the above example,
we can quickly see the boiling point, decomposition point, and melting point of this PFAS.
And literature with undoubtedly even more knowledge about this PFAS:</p>

<p><img src="/assets/images/pfas_scholia.png" alt="" /></p>

<p>Now, these two steps were mostly manual: drawing <a href="https://classic.wikipathways.org/index.php/Pathway:WP5579">WP5579</a>
in WikiPathways and adding the report annotations (<em>main subject</em> and <em>cites</em>) in Wikidata.</p>

<h2 id="findable-in-the-vhp4safety-compound-wiki">Findable in the VHP4Safety Compound Wiki</h2>

<p>As part of the VHP4Safety project, I am collecting information on chemicals studied in the context
of toxicology, safety, and risk assessment. Often specific collections of compounds studied as a whole.
This report is such a collection and provides experimental data on these compounds. So, I want this
report to be findable for the <a href="https://compoundcloud.wikibase.cloud/">VHP4Safety Compound Wiki</a> too.
Creating the collection is a manual step: <a href="https://compoundcloud.wikibase.cloud/wiki/Item:Q5145">Q5145</a>.</p>

<p>Now, because both Wikidata and our VHP4Safety Compound Wiki (a Wikibase instance) are semantic and support, I can use SPARQL
to create instructions to link the 28 compounds to the new collection. Now, arguably, that can be
done manually too, and maybe faster, for larger collections this is harder. So, I dug up
<a href="https://compoundcloud.wikibase.cloud/wiki/User:Egonw">my earlier notes</a> and got some useful
things together.</p>

<p>This query lists all 28 PFAS linked to the report <a href="https://w.wiki/Eepm">in Wikidata</a>:</p>

<div class="language-sparql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span><span class="w"> </span><span class="nv">?pfas</span><span class="w"> </span><span class="nv">?pfasLabel</span><span class="w"> </span><span class="k">WHERE</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q135222054</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P921</span><span class="w"> </span><span class="nv">?pfas</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="nv">?pfas</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P31</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q113145171</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="k">SERVICE</span><span class="w"> </span><span class="nn">wikibase</span><span class="o">:</span><span class="ss">label</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nn">bd</span><span class="o">:</span><span class="ss">serviceParam</span><span class="w"> </span><span class="nn">wikibase</span><span class="o">:</span><span class="ss">language</span><span class="w"> </span><span class="s2">"[AUTO_LANGUAGE],mul,en"</span><span class="p">.</span><span class="w"> </span><span class="p">}</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>Using federation powers, I can use this for <a href="https://edu.nl/ar9wf to match these up with our Wikibase">a SPARQL query</a>,
and return the results in QuickStatements that say <em>this VHP compound is part of the VHP collection</em>:</p>

<div class="language-sparql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">PREFIX</span><span class="w"> </span><span class="nn">wb</span><span class="o">:</span><span class="w"> </span><span class="nn">&lt;https://compoundcloud.wikibase.cloud/entity/&gt;</span><span class="w">
</span><span class="k">PREFIX</span><span class="w"> </span><span class="nn">wbt</span><span class="o">:</span><span class="w"> </span><span class="nn">&lt;https://compoundcloud.wikibase.cloud/prop/direct/&gt;</span><span class="w">

</span><span class="k">SELECT</span><span class="w"> </span><span class="p">(</span><span class="nb">SUBSTR</span><span class="p">(</span><span class="nb">STR</span><span class="p">(</span><span class="nv">?cmp</span><span class="p">),</span><span class="mi">45</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?qid</span><span class="p">)</span><span class="w"> </span><span class="nv">?P21</span><span class="w"> </span><span class="k">WHERE</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="nv">?cmp</span><span class="w"> </span><span class="nn">wbt</span><span class="o">:</span><span class="ss">P5</span><span class="w"> </span><span class="nv">?wikidata</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="k">SERVICE</span><span class="w"> </span><span class="nn">&lt;https://query.wikidata.org/sparql&gt;</span><span class="w"> </span><span class="p">{</span><span class="w">
    </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q135222054</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P921</span><span class="w"> </span><span class="nv">?pfas</span><span class="w"> </span><span class="p">.</span><span class="w">
    </span><span class="nv">?pfas</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P31</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q113145171</span><span class="w"> </span><span class="p">.</span><span class="w">
    </span><span class="k">BIND</span><span class="w"> </span><span class="p">(</span><span class="nb">substr</span><span class="p">(</span><span class="nb">str</span><span class="p">(</span><span class="nv">?pfas</span><span class="p">),</span><span class="mi">32</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?wikidata</span><span class="p">)</span><span class="w">
  </span><span class="p">}</span><span class="w">
  </span><span class="k">BIND</span><span class="w"> </span><span class="p">(</span><span class="s2">"Q5145"</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?P21</span><span class="p">)</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>I actually had to add 5 PFAS compounds in the VHP4Safety Compound Wiki first. That follows the
<a href="https://chem-bla-ics.linkedchemistry.info/2016/03/20/adding-disclosures-to-wikidata-with.html">same procedure for how I have been adding chemical compounds to Wikidata</a>
(see also <a href="https://doi.org/10.26434/chemrxiv-2025-53n0w">this preprint</a>).
The input <code class="language-plaintext highlighter-rouge">cas.smi</code> has the (missing) SMILES, Wikidata QID, and English label:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>C(CS(=O)(=O)O)C(C(C(C(C(C(F)(F)F)(F)F)(F)F)(F)F)(F)F)(F)F       Q27063662       6:2 Fluorotelomer sulfonate
CN(CC(=O)O)S(=O)(=O)C(C(C(C(C(C(F)(F)F)(F)F)(F)F)(F)F)(F)F)(F)F Q126605979      MeFHxSAA
CN(CC(=O)O)S(=O)(=O)C(C(C(C(F)(F)F)(F)F)(F)F)(F)F       Q126682412      MeFBSAA
C(=O)(C(C(F)(F)F)(F)OC(C(C(F)(F)F)(F)F)(F)F)O[H]        Q29387971       2,3,3,3-tetrafluoro-2-(heptafluoropropoxy)propanoic acid
C(C(C(=O)O)(F)F)(OC(C(C(OC(F)(F)F)(F)F)(F)F)(F)F)F      Q81981675       4,8-Dioxa-3H-perfluorononanoic acid
</code></pre></div></div>

<p>For reference, this is the command line I used to create QuickStatement instructions:</p>

<div class="language-shell highlighter-rouge"><div class="highlight"><pre class="highlight"><code>groovy createWDitemsFromSMILES.groovy <span class="nt">-w</span> compoundcloud.wikibase.cloud <span class="nt">-c</span> Q2368 <span class="nt">-d</span> P5 <span class="nt">-l</span> <span class="nt">-i</span> wikidata <span class="nt">-a</span> P11
</code></pre></div></div>

<h2 id="final-remark">Final remark</h2>

<p>Are these 16 the only PFAS in our body? With 28 studied out of <a href="https://doi.org/10.1021/acs.est.3c04855">a potential seven million</a>,
I doubt it.</p>]]></content><author><name>Egon Willighagen</name></author><category term="pfas" /><category term="chemistry" /><category term="fair" /><category term="scholia" /><category term="wikidata" /><category term="vhp4safety" /><category term="doi:10.26434/CHEMRXIV-2025-53N0W" /><category term="cito:citesAsRecommendedReading:10.1021/acs.est.3c04855" /><summary type="html"><![CDATA[A recent report by the Dutch RIVM, PFAS in the blood of the Dutch population (doi:10.21945/RIVM-2025-0094), writes that seven PFAS compounds are found in blood samples of all tested people. Another nine compounds are found in at least 1-in-10 people. Because there is relevant data in the report on the 28 studied PFAS compound, I wanted to have the report more FAIR than it is on the website. Why this report? Well, the chemistry and the history is fascinating and brutal (I like this Veritasium video).]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/pfas_report.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/pfas_report.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">New preprint: “Scholia Chemistry: access to chemistry in Wikidata”</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/05/25/new-preprint-scholia-chemistry-access-to-chemistry-in-wikidata.html" rel="alternate" type="text/html" title="New preprint: “Scholia Chemistry: access to chemistry in Wikidata”" /><published>2025-05-25T00:00:00+00:00</published><updated>2025-05-25T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/05/25/new-preprint-scholia-chemistry-access-to-chemistry-in-wikidata</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/05/25/new-preprint-scholia-chemistry-access-to-chemistry-in-wikidata.html"><![CDATA[<p>Two week ago I uploaded a paper that has been in the works for some time. In fact, I first mention it as conference paper
for the special issue of the <a href="https://scholia.toolforge.org/event/Q47501229">11th International Conference on Chemical Structures</a>,
you know, the meeting held in 2018, of which <a href="https://iccs-nl.org/">the 13th edition</a> starts in 7 days. I had a
<a href="https://doi.org/10.6084/m9.figshare.6356027.v1">poster</a> at that conference which I described in
<a href="https://chem-bla-ics.linkedchemistry.info/2018/08/18/compound-class-identifiers-in-wikidata.html">this blog post</a>.</p>

<p>In turn, that poster described work of at least three years, going back to
<a href="https://chem-bla-ics.linkedchemistry.info/2015/12/22/new-edition-getting-cas-registry.html">adding identifiers in 2015</a>
and <a href="https://chem-bla-ics.linkedchemistry.info/2016/01/27/adding-chemical-compound-to-wikidata.html">chemical structures in early 2016</a>.
I started <a href="https://chem-bla-ics.linkedchemistry.info/2016/03/20/adding-disclosures-to-wikidata-with.html">using scripts two months later</a>.
This helped a lot with <a href="https://chem-bla-ics.linkedchemistry.info/2016/03/27/migrating-pka-data-from-drugmet-to.html">migrating pKa data</a>
from a custom Semantic MediaWiki installation to Wikidata and with adding thousands of EPA CompTox
<a href="https://chem-bla-ics.blogspot.com/2017/01/epa-comptox-dashboard-ids-in-wikidata.html">identifiers in 2017</a>.</p>

<p>But that 2018 conference paper never happened. Because <a href="https://chem-bla-ics.linkedchemistry.info/2017/10/15/two-conference-proceedings.html">Scholia did</a>.
And even on the ICCS poster, Scholia was used to visualize chemistry data in Wikidata. To be honest, not just that,
of course. About a year ago I had a serious go at finishing the paper, and it was sent around to co-authors.
But I realized at the time, that the paper was lacking some good suggestions how the peer review our
actual contributions to Wikidata. I could hardly expect readers of the paper browse the individual
histories of all, by then, 1.3 million chemical compounds. And during the holidays I collected a few
tools, which I had lined up to add to the manuscript.</p>

<p>However, another thing happened, the COVID-19 pandemic. While all the experience helped a lot with getting
knowledge together around SARS-CoV-2, it also made something else clear: the software behind Wikidata
does not scale well (enough). This lead to plans to split the RDF graph representation into two
separate SPARQL endpoints. And that breaks many, if not most, of Scholia’s SPARQL queries, including
those for the chemistry aspects. The situation in Summer 2024 was that there was a significant
chance Scholia would not survive the split. And the <em>Scholia Chemistry</em> paper had to wait. You
cannot publish an article of which the website is gone before it is formally accepted.</p>

<p>Let me make clear, this graph split is not solved and the risk is not gone. But a serious of unfunded,
weekend hackathons allows us to refactor Scholia to give us a chance. It started with
<a href="https://chem-bla-ics.linkedchemistry.info/2024/08/23/scholia.html">making Scholia more configurable</a>.
We had the first hackathons in October and November, and I had
<a href="https://chem-bla-ics.linkedchemistry.info/2025/04/20/the-april-2025-scholia-hackathon.html">four more hackathon weekends</a>
this April.</p>

<p>The graph split into a main graph and a scholarly graph <a href="https://www.wikidata.org/wiki/Wikidata:SPARQL_query_service/WDQS_graph_split">happened on May 9</a>.
Currently, we have been granted extra time and can use a legacy server with the full graph, but a lot
less hardware, so slower. A final patch, merged in last week, allows us to define which SPARQL endpoint a query
should run. So, each time we port a SPARQL query, we can directly update Scholia, making the migration
somewhat more manageable.</p>

<p>But, with those uncertainties out of the way, it was time to finish the Scholia Chemistry paper!</p>

<p>The preprint (doi:<a href="https://doi.org/10.26434/chemrxiv-2025-53n0w">10.26434/chemrxiv-2025-53n0w</a>) brings
10 years of research together, and describes details of the used methods not formally peer-reviewed before.
We describe in detail how chemical structures are added, the choices of Wikidata on how to
represent chemical structures, how we curate the quality, and how we visualize chemical structures
and data with Scholia. As you can expect, the Chemistry Development Kit has an important role,
along with the InChI.</p>

<p>The paper introduces three new Scholia <em>aspects</em> for chemicals, chemical classes, and elements.
Each aspect is a template for a page with information about molecular entities and chemical substances,
compound classes (like <em>fatty acids</em>), and elements (like carbon). Each template provides relevant
information. Of course, any compound, class, or element can also still be opened in the Scholia
“topic” aspect, listing relevant literature.</p>

<p>With this paper we aim to show that Wikidata is a innovative platform that meets the needs for
a chemical structure database, with detailed data provenance, and scalable community curation.</p>

<p>I welcome your strongest peer review on the preprint. I don’t liking settling for anything less.
Here’s the abstract:</p>

<blockquote>
  <p>Sharing knowledge on chemicals in the digital age has been the playground of databases such
as the Chemical Abstract Services and PubChem. Wikipedia complements this field by providing
context to chemicals aimed at a broad audience, but is not easily read by machines. Wikidata
was started as a database service to improve the machine readability of the knowledge captured
in Wikipedia. Wikidata has an open license, application programming interfaces, and a strong
provenance model. Scholia uses the features to provide access to chemical knowledge. This
study reviews the chemistry in Wikidata, shows how thousands of new chemicals were added,
extends Wikidata with new properties for chemical representation and external links to
additional databases, and shows how we extended Scholia to represent the chemistry in Wikidata.</p>
</blockquote>

<p>Thanks to Finn, Denise, Daniel, and Adriano!</p>]]></content><author><name>Egon Willighagen</name></author><category term="wikidata" /><category term="scholia" /><category term="chemistry" /><category term="iccs" /><category term="cito:citesAsEvidence:10.6084/m9.figshare.6356027.v1" /><category term="doi:10.26434/CHEMRXIV-2025-53N0W" /><summary type="html"><![CDATA[Two week ago I uploaded a paper that has been in the works for some time. In fact, I first mention it as conference paper for the special issue of the 11th International Conference on Chemical Structures, you know, the meeting held in 2018, of which the 13th edition starts in 7 days. I had a poster at that conference which I described in this blog post.]]></summary></entry><entry><title type="html">The April 2025 Scholia hackathon</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/04/20/the-april-2025-scholia-hackathon.html" rel="alternate" type="text/html" title="The April 2025 Scholia hackathon" /><published>2025-04-20T00:00:00+00:00</published><updated>2025-04-20T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/04/20/the-april-2025-scholia-hackathon</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/04/20/the-april-2025-scholia-hackathon.html"><![CDATA[<p>This is the third weekend I am working on Scholia, the first two part of the <a href="https://www.wikidata.org/wiki/Wikidata:Scholia/Events/Hackathon_April_2025#Participants">April 2025 hackathon</a>. It follows the hackathons
last year <a href="https://www.wikidata.org/wiki/Wikidata:Scholia/Events/Hackathon_October_2024">October</a> and
<a href="https://chem-bla-ics.linkedchemistry.info/2024/11/17/sparql-examples.html">November</a> hackathons.
There is some urgency for this unpaid work, because Wikidata is splitting the RDF into two
SPARQL endpoints (see <a href="https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2024-05-16/Op-Ed">this The Signpost</a>
and <a href="https://finnaarupnielsen.wordpress.com/2024/10/18/scholia-in-the-age-of-the-wikidata-query-service-split/">this post by Finn</a>).
This split has happened, but there is a <em>legacy</em> server for tools that have not been upgraded.</p>

<p>Scholia has not been upgraded. It has more then 350 SPARQL queries, and each one has to be tested
separately and updating every query is not trivial. Together with Daniel, Finn, and others, I have
hacked up patches last year to:</p>

<ul>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2024/08/23/scholia.html">configure Scholia for the endpoint to use</a></li>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2024/11/17/sparql-examples.html">create pages for many Scholia SPARQL queries</a></li>
</ul>

<p>This month I continued working on the second, and I:</p>

<ul>
  <li><a href="https://github.com/WDscholia/scholia/pull/2597">added titles for chemistry aspect panels</a></li>
  <li><a href="https://github.com/WDscholia/scholia/pull/2589">use the legacy SPARQL endpoint</a></li>
</ul>

<p>That second also indicated that the legacy server has more limited resources and users will more
quickly run into error messages that too many queries are run in parallel. Now, users can rerun
the query, but then the results table contains the previous error message. Second, you want to run
the queries for one aspect to not run all at the same time, but have Scholia send of the
query when the panel becomes visible (and scrolling a page takes a bit of time).</p>

<p>For these issues, I wrote these two patches (yet to be approven and merged):</p>

<ul>
  <li><a href="https://github.com/WDscholia/scholia/pull/2608">delete the previous error message</a></li>
  <li><a href="https://github.com/WDscholia/scholia/pull/2611">lazy load the table and iframe panels</a></li>
</ul>

<p>Now, the iframes already had some aspects of lazy loading, but it turned out that it was mostly
lazy display, and the queries were still run as soon as possible. This last patch challenged my
JavaScript skills and I learned <code class="language-plaintext highlighter-rouge">Intersection Observer API</code>, a browser technology that allows
the browser to see what part of the webpage you are looking at right now. Yeah, I can easily
see how that does user profiling, but in this case it is just used to fire of the SPARQL
query when it become relevant. It uses an additional callback function, so I had to
make sure Jekyll/Liquid creates custom callback functions for each panel.
Actually, I intended to show the code here, but I am not entirely sure how to escape
the code so that Jekyll does not try to run the instructions. For now, you have to
<a href="https://github.com/WDscholia/scholia/pull/2611/files">check the PR</a>.</p>

<h2 id="scholia-chemistry-paper">Scholia Chemistry paper</h2>

<p>The other things I have been doing, is finally finish up the Scholia Chemistry paper.
That actually depended on the maturing of various tools, me figuring out how to characterize
the actual amount of content contributed to Wikidata and how to make that transparent,
and more recently, the above to be able to convince readers Scholia will not die with
the graph split. With the above pages, we have, I think, sufficient guarantee it will
be around for another few years, at least.</p>

<p>This paper, which I hope to finish the final draft today, applying
some good feedback from co-author last weekend, is the final bit of work done on
the Alfred P. Sloan Foundation grant.</p>

<p>We intend to put the paper up as preprint soon and then submit it do a Diamond Open Access
journal, one that supports CiTO citation intent annotation.</p>]]></content><author><name>Egon Willighagen</name></author><category term="scholia" /><category term="javascript" /><category term="sparql" /><summary type="html"><![CDATA[This is the third weekend I am working on Scholia, the first two part of the April 2025 hackathon. It follows the hackathons last year October and November hackathons. There is some urgency for this unpaid work, because Wikidata is splitting the RDF into two SPARQL endpoints (see this The Signpost and this post by Finn). This split has happened, but there is a legacy server for tools that have not been upgraded.]]></summary></entry><entry><title type="html">SPARQL examples: SIB model, software, and patches</title><link href="https://chem-bla-ics.linkedchemistry.info/2024/11/17/sparql-examples.html" rel="alternate" type="text/html" title="SPARQL examples: SIB model, software, and patches" /><published>2024-11-17T00:00:00+00:00</published><updated>2024-11-17T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2024/11/17/sparql-examples</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2024/11/17/sparql-examples.html"><![CDATA[<p><a href="https://akademienl.social/@jerven">Jerven Bolleman</a> <em>et al.</em> recently <a href="https://arxiv.org/abs/2410.06010">published a great preprint</a>
about how to use RDF to give SPARQL queries context by linking it (semantically) with metadata. The context includes
keywords, the SPARQL endpoint the query can be run against, and a human-oriented description of the query. A few groups
have at recent hackathons been working on usingn the combination of a SPARQL query and a human-oriented description
to train large language models, including the group behind this paper. Given that SPARQL is a very small language, I can see
this may work well, and that it may support our <a href="https://vhp4safety.nl/">VHP4Safety</a> and
<a href="https://scholia.toolforge.org/">Scholia</a> projects.</p>

<p>But in addition to the data model for SPARQL as research output (see doi:<a href="https://doi.org/10.32388/ZNWI7T.2">10.32388/ZNWI7T.2</a>),
the paper also introduces the <a href="https://github.com/sib-swiss/sparql-examples-utils">sparql-example-utils</a> software that I was
first introduced with at <a href="https://www.wikidata.org/wiki/Wikidata:Scholia/Events/Hackathon_October_2024">the recent October Scholia hackathon</a>.</p>

<p>But I have/had some features I like to see added. The first is provenance. Who is the author/contributor of the SPARQL
query? Is there a open license for it, or perhaps public domain? How do I give attribution if I reuse the SPARQL query?
These things matter in a modern <a href="https://recognitionrewards.nl/">recognition and rewards</a> world where is room for
everyone’s talent. A set of good SPARQL queries may be more valuable than a ten-page Jupyter notebook (and the other way
around). So, I <a href="https://github.com/sib-swiss/sparql-examples-utils/pull/24">started</a>
<a href="https://github.com/sib-swiss/sparql-examples-utils/pull/25">writing</a>
<a href="https://github.com/sib-swiss/sparql-examples-utils/pull/26">patches</a>. And I created
<a href="https://github.com/BiGCAT-UM/sparql-examples-utils/releases/tag/v2.0.11-tgx-1">a custom jar</a> so that I can see these
patches in action in <a href="https://bigcat-um.github.io/sparql-examples/">our growing list of SPARQL queries</a>
(here <a href="https://bigcat-um.github.io/sparql-examples/examples/WikiPathways/002.html">a WikiPathways query</a>):</p>

<p><img src="/assets/images/sparql-examples-tgx.png" alt="" /></p>

<p>I started collecting SPARQL queries for <a href="https://bigcat-um.github.io/sparql-examples/examples/ChEMBL/">ChEMBL</a>,
<a href="https://bigcat-um.github.io/sparql-examples/examples/WikiPathways/">WikiPathways</a>, and
<a href="https://bigcat-um.github.io/sparql-examples/examples/VHP4Safety/">VHP4Safety</a>. These queries are often part
of other interfaces but we can easily extract the original SPARQL from the Turtle files behind these pages.</p>]]></content><author><name>Egon Willighagen</name></author><category term="sparql" /><category term="doi:10.32388/ZNWI7T" /><category term="justdoi:10.48550/arXiv.2410.06010" /><category term="wikipathways" /><category term="vhp4safety" /><category term="chembl" /><category term="scholia" /><summary type="html"><![CDATA[Jerven Bolleman et al. recently published a great preprint about how to use RDF to give SPARQL queries context by linking it (semantically) with metadata. The context includes keywords, the SPARQL endpoint the query can be run against, and a human-oriented description of the query. A few groups have at recent hackathons been working on usingn the combination of a SPARQL query and a human-oriented description to train large language models, including the group behind this paper. Given that SPARQL is a very small language, I can see this may work well, and that it may support our VHP4Safety and Scholia projects.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/sparql-examples-tgx.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/sparql-examples-tgx.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Scholia configurability</title><link href="https://chem-bla-ics.linkedchemistry.info/2024/08/23/scholia.html" rel="alternate" type="text/html" title="Scholia configurability" /><published>2024-08-23T00:00:00+00:00</published><updated>2024-08-23T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2024/08/23/scholia</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2024/08/23/scholia.html"><![CDATA[<p><a href="https://scholia.toolforge.org/">Scholia</a> is a visual layer on top of <a href="https://wikidata.org/">Wikidata</a> providing
a rich user experience for browing scholarly research related knowledge. I am using the combinatie
for various things, including exploring new research topics (a method, compound, or protein I do not know so much
about yet), indexing notable research output (including citations), <a href="https://chem-bla-ics.linkedchemistry.info/tag/cito">progress of Citation Typing Ontology
uptake</a>, etc. This weekend I hope to send around the
final draft for the <em>Scholia Chemistry</em> paper.</p>

<p>Scholia has received a fair share of scholarly and social attention. The Scholia paper has been cited
<a href="https://scholar.google.com/scholar?hl=en&amp;as_sdt=0%2C5&amp;q=scholia+wikidata&amp;btnG=&amp;oq=scholia">over 100 times</a> and
the websites received about 200 thousand page views each day (though we do not know how to get Toolforge
to give us sufficient insight into the how and what of that count). There is a Wikipedia template to link
to Scholia and some of projects I am involved in link Scholia for articles, such as
<a href="https://wikipathways.org/">WikiPathways</a>.</p>

<p>With that, there is also interest in using it for other Wikibases and perhaps even random SPARQL endpoints.
These things are not trivial, as Scholia uses complementary APIs, various URL patterns for some of the
functionality, and generally, all SPARQL queries are tweaked to the Wikidata Blazegraph SPARQL endpoint
to ensure results are returned in reasonable time. But that last requires use of Blazegraph extensions
to the SPARQL standard.</p>

<p>All this requires Scholia to become more independent, in a better model-view-controller model. And that
actually turns out very important at this moment. That is, Wikidata is not a RDF-first database, but
a Wikibase-based store. Whenever an edit is made, RDF is generated and the SPARQL endpoint is updated.
Now, the number of edits in Wikidata is enormous and the notion that the SPARQL endpoint is often minutes
at most behind is a huge accomplishment. But the Blazegraph platform cannot keep up with Wikidata.
Blazegraph is open source, but has been bought up and development stopped from one day to another.</p>

<p>Therefore, a split of the Wikidata SPARQL platform is <a href="https://phabricator.wikimedia.org/T337013">planned</a>.
This split will put one part of
the knowledge in on endpoint and the other half in the other. Any query that needs information
from both graphs, will have to do a federated SPARQL query. Basically, there are very few Scholia
queries that do not rewriting. My first rewrite actually failed, because the rewriting is not
obvious and quickly times out. To some extend, this is because now lots of results of subqueries
need to be send over the network from one endpoint to the other. When the combined query basically
covers half of each endpoint, that’s a lot of network traffic.</p>

<p>An immediate use case of the configuration is therefore running Scholia against the current three
endpoints: the current official endpoint, and the two split endpoints under development. With
<a href="https://github.com/WDscholia/scholia/pull/2515">a recent patch</a> <a href="@fnielsen@expressional.social">Finn</a>
and I worked on, this configuration looks like this (and saved as <code class="language-plaintext highlighter-rouge">scholia.ini</code>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[query-server]
# Wikidata:
#sparql_endpoint = https://query.wikidata.org/sparql
#sparql_editurl = https://query.wikidata.org/#
#sparql_embedurl = https://query.wikidata.org/embed.html#

# Wikidata Split Main
sparql_endpoint = https://query-main.wikidata.org/sparql
sparql_editurl = https://query-main.wikidata.org/#
sparql_embedurl = https://query-main.wikidata.org/embed.html#

# Wikidata Split Scholar
#sparql_endpoint = https://query-scholarly.wikidata.org/sparql
#sparql_editurl = https://query-scholarly.wikidata.org/#
#sparql_embedurl = https://query-scholarly.wikidata.org/embed.html#
</code></pre></div></div>

<p>So, right now, we can test the impact of the split with Scholia and this patch.
We would fire up a local instances of Scholia, running against one of the
split endpoints, and use the Toolforge instance as baseline.</p>

<p>Now, on my system I need to use <a href="https://python.land/virtual-environments/virtualenv">Python virtualenv</a>
so, I first start a Scholia <code class="language-plaintext highlighter-rouge">venv</code>:</p>

<div class="language-shell highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">source</span> ~/.venvs/scholia/bin/activate
</code></pre></div></div>

<p>After that, I can select an other endpoint, e.g. the <code class="language-plaintext highlighter-rouge">main</code> Wikidata split endpoint (<code class="language-plaintext highlighter-rouge">query-main-experimental.wikidata.org</code>)
were it not they are <a href="https://phabricator.wikimedia.org/T371833">currently offline</a> as part of the transition
and run Scholia on a unique port:</p>

<div class="language-shell highlighter-rouge"><div class="highlight"><pre class="highlight"><code>scholia run
</code></pre></div></div>

<p>Then I can have two browser windows along side and compare Scholia pages againt the current
Scholia instance and when running against another SPARQL endpoint. For now, I can test how well
Scholia runs on the <a href="qlever.cs.uni-freiburg.de/wikidata">QLever instance of Wikidata</a> (superfast and
updated data once a week). Here the configuration I have is not entirely complete, and many
SPARQL queries do not work against QLever, including anything with graphical depiction. But
that said, I can use this configuration:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[query-server]
# QLever
#sparql_endpoint = https://qlever.cs.uni-freiburg.de/api/wikidata
#sparql_editurl = https://qlever.cs.uni-freiburg.de/wikidata/?query=
#sparql_embedurl = 
</code></pre></div></div>

<p>Then, I can compare, for example, the chemicals statistics the main Scholia with one running
against QLever:</p>

<p><img src="/assets/images/scholia_comparison.png" alt="" /></p>

<p>This query ran without modification. For other queries rewriting is needed, but with this
setup we can at least quickly see the differences in the results.</p>]]></content><author><name>Egon Willighagen</name></author><category term="scholia" /><category term="doi:10.1007/978-3-319-70407-4_36" /><category term="wikidata" /><category term="sparql" /><summary type="html"><![CDATA[Scholia is a visual layer on top of Wikidata providing a rich user experience for browing scholarly research related knowledge. I am using the combinatie for various things, including exploring new research topics (a method, compound, or protein I do not know so much about yet), indexing notable research output (including citations), progress of Citation Typing Ontology uptake, etc. This weekend I hope to send around the final draft for the Scholia Chemistry paper.]]></summary></entry><entry><title type="html">New paper: “Wikidata subsetting: approaches, tools, and evaluation”</title><link href="https://chem-bla-ics.linkedchemistry.info/2024/02/13/wikidata-subsetting.html" rel="alternate" type="text/html" title="New paper: “Wikidata subsetting: approaches, tools, and evaluation”" /><published>2024-02-13T00:00:00+00:00</published><updated>2024-02-13T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2024/02/13/wikidata-subsetting</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2024/02/13/wikidata-subsetting.html"><![CDATA[<p>Just before the end of the year, the <em>Wikidata subsetting: approaches, tools, and evaluation</em> paper
by Seyed Amir Hosseini Beghaeiraveri <em>et al.</em> got published (doi:<a href="https://doi.org/10.3233/SW-233491">10.3233/SW-233491</a>).
I am really excited our group (i.e.
<a href="https://orcid.org/0000-0002-8399-8990">Ammar</a> and <a href="https://orcid.org/0000-0001-8449-1318">Denise</a>)
has been able to contribute to this. I think it also is a great example
of the power of hackathons to bring together people.</p>

<p>To me, subsetting of Wikidata (or any large knowledge graph) is important for a couple of reasons.
First, there can be practical reasons. Scholia, for example, is computationally expensive, and the idea
we explore in the Alfred P. Sloan Foundation grant for Scholia (doi:<a href="https://doi.org/10.3897/rio.5.e35820">10.3897/rio.5.e35820</a>)
was that a subset of Wikidata would make it more performant and potentially
more environmental-friendly.</p>

<p>A second reason is more about the scientific process. When doing an analysis and when you want to make
the reasoning transparent, you want to share the analyzed data as part of the research output (basically, the “data”).
For example, the data may have undergone some curation, or you combined data from two or more different
sources. And you will want to share this as part of the scientific process. Resharing a full dump
of the larger knowledge base would not be practical for at least two reasons: duplication of huge data,
and a lot of unrelated content makes it hard for peers to find the bits of interest to the study.</p>

<p>Subsetting may be useful here. This paper evaluates a number of different subsetting approaches.
Myself, I am particularly excited about the idea that we can take a shape expression (e.g. <a href="https://shex.io">ShEx</a>)
as input. I still love the idea that I take the SPARQL queries in my analyses, convert that into
shapes automatically, and then get a subet that returns the exact same results as the query would
on the full dataset.</p>]]></content><author><name>Egon Willighagen</name></author><category term="wikidata" /><category term="doi:10.3233/SW-233491" /><category term="scholia" /><category term="doi:10.3897/RIO.5.E35820" /><summary type="html"><![CDATA[Just before the end of the year, the Wikidata subsetting: approaches, tools, and evaluation paper by Seyed Amir Hosseini Beghaeiraveri et al. got published (doi:10.3233/SW-233491). I am really excited our group (i.e. Ammar and Denise) has been able to contribute to this. I think it also is a great example of the power of hackathons to bring together people.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/wikidata_subsetting_features.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/wikidata_subsetting_features.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>