<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="https://chem-bla-ics.linkedchemistry.info/feed/by_tag/chemometrics.xml" rel="self" type="application/atom+xml" /><link href="https://chem-bla-ics.linkedchemistry.info/" rel="alternate" type="text/html" /><updated>2026-08-31T19:33:33+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/feed/by_tag/chemometrics.xml</id><title type="html">chem-bla-ics</title><subtitle>Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.</subtitle><author><name>Egon Willighagen</name></author><entry><title type="html">AI Technologies in Academia</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/08/18/ai-technologies-in-academia.html" rel="alternate" type="text/html" title="AI Technologies in Academia" /><published>2025-08-18T00:00:00+00:00</published><updated>2025-08-18T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/08/18/ai-technologies-in-academia</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/08/18/ai-technologies-in-academia.html"><![CDATA[<p>I have had the <a href="https://openletter.earth/open-letter-stop-the-uncritical-adoption-of-ai-technologies-in-academia-b65bba1e">Open Letter: Stop the Uncritical Adoption of AI Technologies in Academia</a>
from June 27 open for some time now. I thought I wanted to sign it, but got stuck on the first paragraphs multiple times:</p>

<blockquote>
  <p>With this letter we take a principled stand against the proliferation of so-called ‘AI’ technologies in universities. As an educational institution,
we cannot condone the uncritical use of AI by students, faculty, or leadership. We also call for reconsidering any direct financial relationships
between Dutch universities and AI companies.The unfettered introduction of AI technology leads to contravention of the spirit of the EU Al act. It
undermines our basic pedagogical values and the principles of scientific integrity. It prevents us from maintaining our standards of independence
and transparency. And most concerning, AI use has been shown to hinder learning and deskill critical thought.</p>
</blockquote>

<p>These few lines contain for me more than 25 years of research and I know the complexities. Before I can co-sign this letter,
I need to understand the details. There is no definition of ‘AI’ here and it mentiones the <a href="https://eur-lex.europa.eu/legal-content/NL/TXT/?uri=CELEX:32024R1689">EU AI Act</a>
(I guess, the letter actually writes “Al” (with an <code class="language-plaintext highlighter-rouge">l</code> of letter) act, I notice now after I read the content in another font),
but I have not read the EU AI Act yet (it is 144 pages of legal text).</p>

<h2 id="the-legal-context-of-the-open-letter">The legal context of the Open Letter</h2>

<p>Let me first say, I am not a lawyer (IANAL). I am not versed in the specific legal definitions of tightly defined and controlled
words.</p>

<p>Reading the <em>EU AI Act</em>, I read a reassuring opening statement (repeated later with more context, links to other laws, etc):</p>

<blockquote>
  <p>to promote the uptake of human centric and trustworthy artificial intelligence (AI) while ensuring a high level of protection
of health, safety, fundamental rights as enshrined in the Charter of Fundamental Rights of the European Union (the ‘Charter’),
including democracy, the rule of law and environmental protection, to protect against the harmful effects of AI systems in the Union</p>
</blockquote>

<p>We clearly see how these things are currently routinely violated.</p>

<blockquote>
  <p>This Regulation does not apply to AI systems or AI models, including their output, specifically developed and put into service
for the sole purpose of scientific research and development.</p>
</blockquote>

<p>In Dutch this is officially translated to “wetenschappelijk onderzoek”, so <em>scientific research</em> seems to be legally
including humanites, etc, and not limited to natural sciences [citation needed].</p>

<p>The EU AI Act also outlines a definition of “AI”, leaning towards machine learning, but the border between deterministic,
rule-based algorithms and machine-learned patters for predictions remains a bit vague to me. But I can live with it.</p>

<p>The Open Letter’s <em>contravention of the spirit of the EU Al act</em> gets context here too. It has to be the <em>spirit</em>,
because the law does not apply to academia. Good, clarified. The Letter continues with:</p>

<blockquote>
  <p>It undermines our basic pedagogical values and the principles of scientific integrity.
It prevents us from maintaining our standards of independence and transparency. And most concerning, AI use has been
shown to hinder learning and deskill critical thought.</p>
</blockquote>

<p>Yes, that clearly links to the EU AI Act’s protection of rights. Maybe on purpose and maybe there are legal reasons
to not explicitly list them, are the international human rights, which includes rigths to benefit from science,
but I think this is still in the spirit of the EU AI Act. And if AI fetters our ability to learn (yes, there
is scientific evidence for that [citation needed]), then it violates the EU AI Act (IANAL).</p>

<h2 id="what-the-open-letter-expects">What the Open Letter expects</h2>

<p>The next part of the Open Letter calls to what the signers expect from our universities. I will will reflect on each of them.</p>

<blockquote>
  <p><strong>Resist the introduction of AI in our own software systems</strong>, from Microsoft to OpenAI to Apple. It is not in our interests
to let our processes be corrupted and give away our data to be used to train models that are not only useless to us, but
also harmful.</p>
</blockquote>

<p>The intrinsic problem and why I think it is fair to call out these companies, is, as the letter explains, there
is an clear conflict of interest. The goal of companies is to make profit (and in a Western world, as much
as possible), and not any of the human or scientific needs. In this respect, companies like Elsevier
could just as well have mentioned too (see e.g. <a href="https://irisvanrooijcogsci.com/2025/08/12/ai-slop-and-the-destruction-of-knowledge/">this post by Prof. Van Rooij</a>,
actually 2nd signature on the letter).</p>

<blockquote>
  <p><strong>Ban AI use in the classroom</strong> for student assignments, in the same way we ban essay mills and other forms
of plagiarism. Students must be protected from de-skilling and allowed space and time to perform their
assignments themselves.</p>
</blockquote>

<p>About a year ago, I was pleasently surprised by the depth of discussion at Maastricht University on how and when
to use AI, and by default not. This one is really complicated and it matters when and how the AI is used.
After all, and the spirit of the EU AI Act expects us to use AI in research (to trigger innovation). So,
I cannot agree with the literal statement, but I fully agree with the spirit. Particularly combined with
the clear “Stop the Uncritical Adoption of AI Technologies in Academia” of the title of the Open Letter.</p>

<p>I read this line like this, AI in the classroom must have a purpose that aligns with the EU AI Act.
That means, use for writing assays, reports, it must not be used. I am old enough that remember the
academic discussions (at Radboud University) about writing and the clear hesitance among scholars
about the use of written assignments: “I want to test their scientific knowledge and reasoning skills,
not their ability to write narratives”. And LLMs, like ChatGPT but also the European, more open variants,
they write narratives, so the written report and assay is no longer a valid way to assess a student’s
scientific learning progress.</p>

<p>So, alternatively, we should very carefully and scientifically evaluate which forms of assessment
we perform, and banning AI in the classroom may just be distracting from a more fundamental problem.
Anyways… if you continue using writing assignments to test progress in learning, you must ban
use of AI in that process. You must be testing the student, not some piece of software (as a teaching
institute).</p>

<blockquote>
  <p><strong>Cease normalising the AI hype</strong> and the lies which are prevalent in the technology industry’s framing of
these technologies. The technologies do not have the advertised capacities and their adoption puts students
and academics at risk of violating ethical, legal, scholarly, and scientific standards of reliability,
sustainability, and safety.</p>
</blockquote>

<p>Sounds like a no brainer. But I too find my own university uncritically promoting AI. Maybe the tested
it well, and just forgot to share that. But hey, scientifical quality goes all ways.</p>

<blockquote>
  <p><strong>Fortify our academic freedom</strong> as university staff to enforce these principles and standards in our
classrooms and our research as well as on the computer systems we are obliged to use as part of our
work. We as academics have the right to our own spaces.</p>
</blockquote>

<p>Again, a no brainer. But important to add. It must be said as it is intrisic part of
<a href="https://recognitionrewards.nl/">Recognition &amp; Rewards</a>. If you cannot guarantee academic freedom,
there there is something seriously wrong with your R&amp;R.</p>

<blockquote>
  <p><strong>Sustain critical thinking on AI</strong> and promote critical engagement with technology on a firm
academic footing. Scholarly discussion must be free from the conflicts of interest caused by
industry funding, and reasoned resistance must always be an option.</p>
</blockquote>

<p>Yeah, this is something that is underestimated. Part of our academic teaching is this critical thinking.
It returns in academic reading (did you already read <em>“What Little Red Riding Hood Can Teach Us about Reading Science”</em>,
doi:<a href="https://uplopen.com/chapters/e/10.1515/9783110782844-010">10.1515/9783110782844-010</a>,
by <a href="https://scholar.google.com/citations?user=0KRmIbcAAAAJ&amp;hl=nl&amp;oi=ao">Monica Gonzalez-Marquez</a> <em>et al.</em>?),
scientific programming, data analysis, and our teaching has been
lacking here. Not just for new AI forms, but also for the old algorihmts. I have seen this, and
scientific literature is riddled with mistakes, just because our peer reviewers are not sufficiently
skilled. This will take effort. I know, it was a major part of
<a href="https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work.html">my PhD thesis</a>.</p>

<p>Of course, this is exactly why I have been so active in Open Science. Without Open Science,
we cannot work <em>in the spirit</em> of the EU AI Act. It’s nothing new. It’s just that the big money
has found in AI a way to profit at the expense of humans.</p>

<p>So, go read that <a href="https://openletter.earth/open-letter-stop-the-uncritical-adoption-of-ai-technologies-in-academia-b65bba1e">Open Letter</a>
and sign too!</p>]]></content><author><name>Egon Willighagen</name></author><category term="cheminf" /><category term="chemometrics" /><category term="justdoi:10.1515/9783110782844-010" /><category term="openscience" /><summary type="html"><![CDATA[I have had the Open Letter: Stop the Uncritical Adoption of AI Technologies in Academia from June 27 open for some time now. I thought I wanted to sign it, but got stuck on the first paragraphs multiple times:]]></summary></entry><entry><title type="html">The Molecular Chemometrics Principles #2: be clear in what you mean</title><link href="https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be.html" rel="alternate" type="text/html" title="The Molecular Chemometrics Principles #2: be clear in what you mean" /><published>2010-08-12T00:00:00+00:00</published><updated>2010-08-12T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be.html"><![CDATA[<p>I noted <a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">earlier this week</a>
that <em>[d]uring the week [in <a href="/2010/08/06/oxford-2.html">Oxford <i class="fa-solid fa-recycle fa-xs"></i></a>], someone (name and address is know at the
editorial office) commented on the fact that my blog posts are somewhat difficult to follow; that is, it’s
often not clear why I am posting what I am posting</em>. This triggered the start of a series of principles in
the field I coined <a href="https://doi.org/10.1080/10408340600969601">Molecular Chemometrics</a>, and the promise
that I will try to indicate in each blog post to which of these principles it relates. Just to put things in a bit more
perspective; to make a bit more clear why I am blogging about that bit; just to be clear in what I mean.</p>

<p>Now, the first principle was about the need for access to data (<a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">McPrinciple #1</a>).
This principle goes without saying, one would think, but is not widely accepted yet. This is why Open Data promotion is still needed. For example, data in papers
still is not freely redistributable, as <a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">Peter points out once again</a>.</p>

<p>Anyway, this post is not about McPrinciple #1, but about the second principle.</p>

<p><strong>Molecular Chemometrics Principles #2</strong>: In order to reproduce cheminformatics studies you need to be able to understand the input data.</p>

<p>Readers of my blog will surely recognize this theme. Clearly this theme explains my past fetish for the
<a href="http://chem-bla-ics.blogspot.com/search?q=CML">Chemical Markup Language</a>, and my more recent work on the
<a href="http://chem-bla-ics.blogspot.com/search?q=RDF">Resource Description Framework</a>.</p>

<p>And it is so easy to jump to conclusions. Easy to make mistakes. And this is not just at the received side; the sending
person may have accidentally made a mistake, or left something accidentally unclear, causing incorrect assumptions, and
therefore errors in the cheminformatics computation. Now, if the data was semantically (clearly) annotated, and the
meaning was clear, it was also trivial to see when a mistake had sneaked in. Think of it as a check bit.</p>

<p>“Well, isn’t this a bit exaggerated,” you might say. Perhaps, perhaps not. An simple, recent example. We all know
<a href="http://www.opensmiles.org/">SMILES</a>, right? And we all know that lower case element symbols indicate aromaticity, right?
That is, c1ccccc1 is aromatic, right? So, what’s the problem then?</p>

<p>Now, consider the SMILES string c1ccc1. Lower case carbon element symbols, so aromatic, right? Oh, wait…</p>

<p>Therefore, be clear in what you mean. It saves us from a lot of trouble.</p>

<p>Further reading:</p>

<ul>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">The Molecular Chemometrics Principles #1: access to data</a></li>
  <li>Molecular Chemometrics, 2006 (doi:<a href="https://doi.org/10.1080/10408340600969601">10.1080/10408340600969601</a>)</li>
</ul>]]></content><author><name>Egon Willighagen</name></author><category term="mcprinciples" /><category term="chemometrics" /><category term="rdf" /><category term="cml" /><category term="semweb" /><category term="doi:10.1080/10408340600969601" /><summary type="html"><![CDATA[I noted earlier this week that [d]uring the week [in Oxford ], someone (name and address is know at the editorial office) commented on the fact that my blog posts are somewhat difficult to follow; that is, it’s often not clear why I am posting what I am posting. This triggered the start of a series of principles in the field I coined Molecular Chemometrics, and the promise that I will try to indicate in each blog post to which of these principles it relates. Just to put things in a bit more perspective; to make a bit more clear why I am blogging about that bit; just to be clear in what I mean.]]></summary></entry><entry><title type="html">The Molecular Chemometrics Principles #1: access to data</title><link href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html" rel="alternate" type="text/html" title="The Molecular Chemometrics Principles #1: access to data" /><published>2010-08-09T00:00:00+00:00</published><updated>2010-08-09T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html"><![CDATA[<p>The meetings in and around Oxford were great! I already wrote that the Predictive Toxicology workshop was brilliant
(see <a href="/2010/08/01/oxford.html">Oxford… #1 <i class="fa-solid fa-recycle fa-xs"></i></a>) and
<a href="/2010/08/06/oxford-2.html">Oxford… #2 <i class="fa-solid fa-recycle fa-xs"></i></a>), but I also very, very much enjoyed meeting up
with <a href="http://www.danhagon.me.uk/blog/">Dan</a> and <a href="http://semanticscience.wordpress.com/">Nico</a>! During the week, someone
(name and address is know at the editorial office) commented on the fact that my blog posts are somewhat difficult
to follow; that is, it’s often not clear why I am posting what I am posting.</p>

<p>Indeed, I am not particularly one of those bloggers who spends trees after trees, in great detail explaining what is going on.
I do make a lot of use of <a href="http://en.wikipedia.org/wiki/Hyperlink">hyperlinking</a>; much more than the average blogger. I
actually assume that readers follow links, to read about the perspective of a blog post. But we all know that scientists
do not read the cited papers in a paper they are reading, so who am I to assume blog readers would start doing that with blogs :)</p>

<p>Well, since <a href="/2010/02/19/open-data-panton-principles.html">principles seems popular <i class="fa-solid fa-recycle fa-xs"></i></a>, it might be
a good start of my grand scheme that is behind this blog: the Molecular Chemometrics Principles. Hence, this first post about
the why. The why is simply to provide a reference frame to what I am blogging about. In the next few posts on these
McPrinciples (is that a catchy name, or what?) that will appear over the next two weeks, I will outline the code of
chem-bla-ics. And, moreover, from now on, I will tag all my posts with the reaons why I make that post. I am sure that will
not be too helpful for the occasional reader, but for anyone who is serious about chem-bla-ics, this will be a genuine gold
mine of data for pattern recognition and data mining otherwise.</p>

<p>So, here goes.</p>

<p><strong>Molecular Chemometrics Principles #1</strong>: In order to reproduce cheminformatics studies you need access to the input data.</p>

<p>The reason for this is that statistical modeling very much depends on the data on which modeling was done, patterns
were recognized, etc. Therefore, without the input data, it is practically impossible to accurately reproduce results.
Fortunately, the acceptance of the importance of access to data (e.g. as Open Data) is slowly getting momentum in
science.</p>

<p>Further reading: Molecular Chemometrics, 2006 (doi:<a href="https://doi.org/10.1080/10408340600969601">10.1080/10408340600969601</a>)</p>]]></content><author><name>Egon Willighagen</name></author><category term="chemometrics" /><category term="mcprinciples" /><category term="doi:10.1080/10408340600969601" /><summary type="html"><![CDATA[The meetings in and around Oxford were great! I already wrote that the Predictive Toxicology workshop was brilliant (see Oxford… #1 ) and Oxford… #2 ), but I also very, very much enjoyed meeting up with Dan and Nico! During the week, someone (name and address is know at the editorial office) commented on the fact that my blog posts are somewhat difficult to follow; that is, it’s often not clear why I am posting what I am posting.]]></summary></entry><entry><title type="html">Looking at your statistical models…</title><link href="https://chem-bla-ics.linkedchemistry.info/2010/06/20/looking-at-your-statistical-models.html" rel="alternate" type="text/html" title="Looking at your statistical models…" /><published>2010-06-20T00:10:00+00:00</published><updated>2010-06-20T00:10:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2010/06/20/looking-at-your-statistical-models</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2010/06/20/looking-at-your-statistical-models.html"><![CDATA[<p>I do not think I have ever blogged the paper that played an important role in my thesis (doi:<a href="http://dx.doi.org/10.1021/ci990038z">10.1021/ci990038z</a>);
research of one of the papers in my thesis, started with the hypothesis proposed therein. The paper had a really good idea; but, unfortunately, it did
not contain the data to support the hypothesis. That gets me to one important lesson I learned: <strong>a QSAR data set of less than 100 molecules is not enough
to make untargeted statistical models.</strong></p>

<p>The paper reads quite nicely, and the results are clear: by combining spectral types, the RMSEP goes down. Good! Lower prediction errors;
that’s what we all want. So, a M.Sc. student of mine set off, but after about half a year, he was still unable to make statistically good models.
He used bootstrapping to ‘prove’ it was not his fault: there was not enough data for the method to learn the underlying patters. Hence my above lesson.
My student went on with larger data sets, and laid out the foundation of what later became the paper on using NMR spectra in QSPR modeling
(doi:<a href="https://doi.org/10.1021/ci050282s">10.1021/ci050282s</a>). Now you understand why QSAR is missing.</p>

<p>So, if those results are so clear, then why does it not work? As said, the data set was too small for pattern recognition methods to see what was going
on. The RMSEP numbers just came out nicely; however, if we had only made the below plot, if would have warned. But I failed to do that at the time. Lesson
learned: <strong>do not just look at the data, but also look at the model.</strong> And look really means looking with your eyes at graphical representations of that model.
The plot:</p>

<p><img src="/assets/images/badModel.png" alt="" /></p>

<p>The numbers in this plot are hidden in tables in the paper. The RMSEP values earlier mentioned are calculated from those. From the plot, you can see
that the test data consisted of 5 compounds; the training set contained 37 compounds; all are congenerics, and do not span a high diversity. Now, the
plot shows five models: black is COMFA; orange is based on experimental IR spectra; red, green, and blue are models where two types of representations
are combined. From the RMSEP values it can be seen that combining representation improves the RMSEP values. That’s what you want, and sort of makes
sense.</p>

<p>Now, I did not make this plot until I started writing up the paper, and tried to figure out why the QSAR data set did not work. My eyes opened wide when
I saw the orange dots! Anti-correlation! WTF?!?! I mean, we are looking at a plot visualizing the predicted versus the experimental activity… Actually,
the others are not really convincing either, are they? Looking at the predictions for compounds with experimental values around 1.0-1.5 (if you really
want to know the unit, read the article), the pattern is pretty much anti-correlated too. Thinking about it, it seems the RMSEP is mainly reflecting the
error of the left most compound, the one with a experimental activity of about 0.3.</p>

<p>Clearly, the orange model is hopeless, but the others are not really better. Now, the paper actually makes statements comparing the various combinations
of representations, but, in retrospect and looking at this plot, I wonder if the green model is really different from the blue or red models.</p>

<p>Since then, I always make these kind of plots, just to see what my model is like. Since then, I distrust papers that only show RMSEP, Q², or other
quality statistics. Now, the tricky part is, you need those statistics if you want to automate model selection; the variance on those model quality
statistics is actually so high (see also <a href="http://chem-bla-ics.blogspot.com/2010/06/qspr-modeling-with-signatures.html">my other post today</a>),
that you must carefully validate that model selection too, visually of course.</p>

<p>I have been long thinking what to do with these observations. I did not dare publish them in my thesis; I did not dare write a letter to the editor.
Perhaps I should. But even writing up this blog makes me feel uncomfortable. Besides the fact that I might be wrong, I also do not like to point out
mistakes (IMHO); particularly, when those are published in a respectable journal. I was fooled by the statistics too (and was already well trained),
so I cannot comment on the authors overlooking the issue. Or the reviewers! Or the community at large. Also, I do not know what the fate of this
paper should be. The idea is quite interesting, even though the published results do not support it. Not shown here, but the bootstrapping results
show that the apparent slight improvement is merely a numerical artifact, just happening by chance, based on luckily selecting the test compounds;
the data is just insufficient in size to draw any conclusion.</p>]]></content><author><name>Egon Willighagen</name></author><category term="chemometrics" /><category term="justdoi:10.1021/ci990038z" /><category term="doi:10.1021/CI050282S" /><category term="qsar" /><summary type="html"><![CDATA[I do not think I have ever blogged the paper that played an important role in my thesis (doi:10.1021/ci990038z); research of one of the papers in my thesis, started with the hypothesis proposed therein. The paper had a really good idea; but, unfortunately, it did not contain the data to support the hypothesis. That gets me to one important lesson I learned: a QSAR data set of less than 100 molecules is not enough to make untargeted statistical models.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/badModel.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/badModel.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">QSPR modeling with signatures</title><link href="https://chem-bla-ics.linkedchemistry.info/2010/06/20/qspr-modeling-with-signatures.html" rel="alternate" type="text/html" title="QSPR modeling with signatures" /><published>2010-06-20T00:00:00+00:00</published><updated>2010-06-20T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2010/06/20/qspr-modeling-with-signatures</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2010/06/20/qspr-modeling-with-signatures.html"><![CDATA[<p>I had to dig deep to find posts on QSAR modeling. There are quite a few on <a href="http://chem-bla-ics.blogspot.com/search?q=qsar+bioclipse">QSAR in Bioclipse</a>,
but that focuses on the descriptor calculation. In a quick scan, I could only spot two modeling posts:</p>

<ul>
  <li><a href="http://chem-bla-ics.blogspot.com/2008/04/cdkmetabolomicschemometrics.html">The CDK/Metabolomics/Chemometrics Unconference results</a></li>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2005/11/08/when-to-stop-including-qsar-model.html">When to stop including QSAR model variables… <i class="fa-solid fa-recycle fa-xs"></i></a></li>
</ul>

<p>Given the prominent place <a href="https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work.html">QSAR has in my thesis <i class="fa-solid fa-recycle fa-xs"></i></a>,
this is somewhat surprising. Anyway, here is some more QSAR modeling talk.</p>

<p><a href="http://gilleain.blogspot.com/">Gilleain</a> <a href="http://www.blogger.com/github.com/gilleain/signatures">implemented</a> the signature descriptors developed by Faulon et al.
(see doi:<a href="https://doi.org/10.1021/ci020345w">10.1021/ci020345w</a>; I <a href="http://chem-bla-ics.blogspot.com/2006/02/novel-qsar-and-qspr-descriptors_24.html">mentioned the paper in 2006</a>),
and the <a href="http://sourceforge.net/tracker/?func=detail&amp;aid=3017759&amp;group_id=20024&amp;atid=320024">CDK patch</a> is currently being reviewed.
With some transformations, the atomic signatures for a molecule can be transformed into a fixed-length numerical representation:
<code class="language-plaintext highlighter-rouge">[70:1, 54:1, 23:1, 22:1, 9:9, 45:2]</code>. This string means that atomic signature 70 occurs once in this molecule and signature 9 occurs
nine times. At this moment, I am not yet concerned about the actual signature, but just checking how well these signature can be used
in QSPR modeling.</p>

<p><a href="http://blog.rguha.net/">Rajarshi</a>’s <a href="http://cran.r-project.org/web/packages/fingerprint/index.html">fingerprint</a> code provides a good
template to parse this into a X matrix in <a href="http://www.r-project.org/">R</a>:</p>

<p>For my test case, I have used the boiling point data I used in my thesis paper <em>On the Use of 1H and 13C 1D NMR Spectra as QSPR Descriptors</em>
(see doi:<a href="https://doi.org/10.1021/ci050282s">10.1021/ci050282s</a>). Some of this data is actually <a href="http://www.chemspider.com/blog/gathering-physicochemical-data-onto-chemspider.html">available from ChemSpider</a>,
but I do not think I ever uploaded the boiling point data. This constitutes a data set with 277 molecules, and my paper provides
some reference model quality statistics; that way, I have something to compare against. Moreover, I can use my previous scripts
to do the PLS modeling (there are many <a href="http://www.google.se/search?q=tutorial+partial+least+squares">tutorials online</a>, but you
can always buy an expensive book like the one shown on the right, if you really have to), (10-fold) cross-validation (CV), and
5 repeats of random sampling.</p>

<p>I strongly suggest people interested in statistical modeling to read this
<a href="http://baoilleach.blogspot.com/2010/06/non-random-method-to-improve-your-qsar.html">interesting post from Noel</a>: whatever test
set sampling method you use, you <strong><em>must</em></strong> do some repeats to learn about the sensitivity of your modeling approach to changes
in the data set. Depending on the actual sampling approach, you might see different sizes of variance, but until you measure it,
you will not know. For my application, these are the numbers:</p>

<div class="language-R highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># source("pls.R")</span><span class="w">
</span><span class="n">Read</span><span class="w"> </span><span class="m">277</span><span class="w"> </span><span class="n">items</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">66</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.987</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.921</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">31.37</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">66</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.983</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.924</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">12.405</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">65</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.985</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.949</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">38.503</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">63</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.983</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.948</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">36.981</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">65</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.986</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.923</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">21.49</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">64</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.983</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.91</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">17.759</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">64</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.983</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.921</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">17.062</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">66</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.986</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.94</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">40.311</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">66</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.982</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.927</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">13</span><span class="w">
   </span><span class="n">rank</span><span class="o">=</span><span class="m">68</span><span class="w">     </span><span class="n">LV</span><span class="o">=</span><span class="m">42</span><span class="w">     </span><span class="n">R2</span><span class="o">=</span><span class="m">0.986</span><span class="w">     </span><span class="n">Q2</span><span class="o">=</span><span class="m">0.929</span><span class="w">  </span><span class="n">RMSEP</span><span class="o">=</span><span class="m">16.23</span><span class="w">
</span></code></pre></div></div>

<p>I know 42 is the answer to the universe, but 42 latent variables (LVs)?!? Well, it’s just a start. A more accurate number of LVs
seems to be around 15, but my script had to make the transition from the old pls.pcr package to the newer pls package. And I have
yet to discover how I can get the new package to return me the lowest number of LVs for which the CV statistic is no longer
significantly different from the best (see my paper how that works). Actually, I have set the maximum LVs to consider to 1/5th of
the number of objects (which is about the accepted ratio in the QSAR community); otherwise, it would have happily increased.</p>

<p>However, the five repeats nicely show the variance in the quality statistics, R², Q², and root mean square error of prediction
(RMSEP). From the numbers, a model with Q² = 0.94 is <strong>not</strong> better than one with Q² = 0.93 (and I have seen the variance quite some
larger). Bottom line: just measure that variability, and put it in the publication, will you??</p>

<p>Anyway, what we all have been waiting for: the prediction results visualized (in black the CV predictions; in red the test set
predictions):</p>

<p><img src="/assets/images/signaturePrediction.png" alt="" /></p>

<p>Well, there is still much work to do, and you can expect the result to get better. Part of statistical modeling is to find
the source of variance, and I have yet to explore a few of them. For example, what are the effects of:</p>

<ul>
  <li>creating signature from the hydrogen-depleted graph</li>
  <li>effect of tautomerism (see <a href="http://www.springerlink.com/content/l3p3t7066645/?p=bff6cd9b91bd40c59aa0d7afe11cf78a&amp;pi=0">this special issue</a>)</li>
  <li>effect of the height of the signature</li>
</ul>

<p>And there are so many other things I like to do. But this will do for now.</p>]]></content><author><name>Egon Willighagen</name></author><category term="cdk" /><category term="chemometrics" /><category term="justdoi:10.1021/ci020345w" /><category term="doi:10.1021/ci050282s" /><summary type="html"><![CDATA[I had to dig deep to find posts on QSAR modeling. There are quite a few on QSAR in Bioclipse, but that focuses on the descriptor calculation. In a quick scan, I could only spot two modeling posts:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/signaturePrediction.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/signaturePrediction.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Moved to Sweden: Post-doc in the Bioclipse group of Prof. Jarl Wikberg</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/09/24/moved-to-sweden-post-doc-in-bioclipse.html" rel="alternate" type="text/html" title="Moved to Sweden: Post-doc in the Bioclipse group of Prof. Jarl Wikberg" /><published>2008-09-24T00:00:00+00:00</published><updated>2008-09-24T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/09/24/moved-to-sweden-post-doc-in-bioclipse</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/09/24/moved-to-sweden-post-doc-in-bioclipse.html"><![CDATA[<p>The reason why I have not been able to blog much lately, is that my family and I have been moving to
<a href="http://en.wikipedia.org/wiki/Uppsala">Uppsala</a>/Sweden, where I’ll start a postdoc in the
<a href="http://www.farmbio.uu.se/researchgroup.php?fg=1">group of Jarl Wikberg</a> @ <a href="http://www.bmc.uu.se/">BMC</a> @
<a href="http://en.wikipedia.org/wiki/Uppsala_University">Uppsala University</a>, where I’ll work on chemoinformatics
in drug design, and the use of <a href="http://cdk.sf.net/">CDK</a> and <a href="http://www.bioclipse.net/">Bioclipse</a>
in particular.</p>

<p>More blogging when I have more frequent internet access again…</p>]]></content><author><name>Egon Willighagen</name></author><category term="cdk" /><category term="cheminf" /><category term="chemometrics" /><category term="career" /><category term="bioclipse" /><summary type="html"><![CDATA[The reason why I have not been able to blog much lately, is that my family and I have been moving to Uppsala/Sweden, where I’ll start a postdoc in the group of Jarl Wikberg @ BMC @ Uppsala University, where I’ll work on chemoinformatics in drug design, and the use of CDK and Bioclipse in particular.]]></summary></entry><entry><title type="html">The CDK/Metabolomics/Chemometrics Unconference results</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/04/07/cdkmetabolomicschemometrics.html" rel="alternate" type="text/html" title="The CDK/Metabolomics/Chemometrics Unconference results" /><published>2008-04-07T00:10:00+00:00</published><updated>2008-04-07T00:10:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/04/07/cdkmetabolomicschemometrics</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/04/07/cdkmetabolomicschemometrics.html"><![CDATA[<p>As <a href="https://chem-bla-ics.linkedchemistry.info/2008/04/03/t-plus-18-hours-dr-and-preparing-for.html">announced earlier <i class="fa-solid fa-recycle fa-xs"></i></a>, Miguel, Velitchka,
<a href="http://www.steinbeck-molecular.de/steinblog/">Christoph</a> and I held a small <a href="http://cdk.sf.net/">CDK</a>/Metabolomics/Chemometrics
unconference. We started late, and did not have an evening program, resulting in not overly much results. However, we did do
<em><a href="http://chem-bla-ics.blogspot.com/search?q=molecular+chemometrics">molecular chemometrics</a></em>. <!-- keep link --></p>

<p>We used the <a href="http://www.r-project.org/">R statistics software</a> together with Rajarshi’s <a href="http://cran.r-project.org/web/packages/rcdk/index.html">rcdk</a>
package (an R wrapper around the CDK library) and Ron’s (my PhD supervisor) <a href="http://cran.r-project.org/web/packages/pls/index.html">PLS</a>
package (see <a href="http://www.jstatsoft.org/v18/i02/">this paper</a>), to predict retention indices for a number of metabolites.</p>

<p>We ended up with this R script:</p>

<div class="language-R highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">library</span><span class="p">(</span><span class="s2">"rJava"</span><span class="p">)</span><span class="w">
</span><span class="n">library</span><span class="p">(</span><span class="s2">"rcdk"</span><span class="p">)</span><span class="w">
</span><span class="n">library</span><span class="p">(</span><span class="s2">"pls"</span><span class="p">)</span><span class="w">
</span><span class="n">mols</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">load.molecules</span><span class="p">(</span><span class="s2">"data_cdk.sdf"</span><span class="p">)</span><span class="w">
</span><span class="n">selection</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">get.desc.names</span><span class="p">()</span><span class="w">
</span><span class="n">selection</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">selection</span><span class="p">[</span><span class="o">-</span><span class="n">which</span><span class="p">(</span><span class="n">selection</span><span class="o">==</span><span class="s2">"org.openscience.cdk.qsar.descriptors.molecular.AminoAcidCountDescriptor"</span><span class="p">)]</span><span class="w">
</span><span class="n">x</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">eval.desc</span><span class="p">(</span><span class="n">mols</span><span class="p">,</span><span class="w"> </span><span class="n">selection</span><span class="p">,</span><span class="w"> </span><span class="n">verbose</span><span class="o">=</span><span class="kc">TRUE</span><span class="p">)</span><span class="w">
</span><span class="n">x2</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">x</span><span class="p">[,</span><span class="n">apply</span><span class="p">(</span><span class="n">x</span><span class="p">,</span><span class="w"> </span><span class="m">2</span><span class="p">,</span><span class="w"> </span><span class="k">function</span><span class="p">(</span><span class="n">a</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="nf">all</span><span class="p">(</span><span class="o">!</span><span class="nf">is.na</span><span class="p">(</span><span class="n">a</span><span class="p">))})]</span><span class="w">
</span><span class="n">y</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">read.table</span><span class="p">(</span><span class="s2">"data_cdk_RI"</span><span class="p">)</span><span class="w">
</span><span class="n">input</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">data.frame</span><span class="p">(</span><span class="n">x2</span><span class="p">,</span><span class="w"> </span><span class="n">y</span><span class="p">)</span><span class="w">
</span><span class="n">pls.model</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">plsr</span><span class="p">(</span><span class="n">V1</span><span class="w"> </span><span class="o">~</span><span class="w"> </span><span class="n">.</span><span class="p">,</span><span class="w"> </span><span class="m">50</span><span class="p">,</span><span class="w"> </span><span class="n">data</span><span class="o">=</span><span class="n">input</span><span class="p">,</span><span class="w"> </span><span class="n">validation</span><span class="o">=</span><span class="s2">"CV"</span><span class="p">)</span><span class="w">
</span><span class="n">summary</span><span class="p">(</span><span class="n">pls.model</span><span class="p">)</span><span class="w">
</span><span class="n">plot</span><span class="p">(</span><span class="n">RMSEP</span><span class="p">(</span><span class="n">pls.model</span><span class="p">))</span><span class="w">
</span><span class="n">plot</span><span class="p">(</span><span class="n">pls.model</span><span class="p">,</span><span class="w"> </span><span class="n">ncomp</span><span class="o">=</span><span class="m">20</span><span class="p">)</span><span class="w">
</span><span class="n">abline</span><span class="p">(</span><span class="m">0</span><span class="p">,</span><span class="m">1</span><span class="p">,</span><span class="w"> </span><span class="n">col</span><span class="o">=</span><span class="s2">"red"</span><span class="p">)</span><span class="w">
</span><span class="n">plot</span><span class="p">(</span><span class="n">pls.model</span><span class="p">,</span><span class="w"> </span><span class="s2">"loadings"</span><span class="p">,</span><span class="w"> </span><span class="n">comps</span><span class="o">=</span><span class="m">1</span><span class="o">:</span><span class="m">2</span><span class="p">)</span><span class="w">
</span><span class="n">savehistory</span><span class="p">(</span><span class="s2">"finalHistory.R"</span><span class="p">)</span><span class="w">
</span></code></pre></div></div>

<p>The <code class="language-plaintext highlighter-rouge">AminoAcidCountDescriptor</code> threw us a <code class="language-plaintext highlighter-rouge">NullPointerException</code> and there were a few NAs in the resulting matrix. The CV results were
not so good as Velitchka’s best models, but still a good start:</p>

<p><img src="/assets/images/riPred.png" alt="" /></p>

<p>No variable selection; 200 objects, 190 variables.</p>

<p>Questions:</p>

<ul>
  <li>Can we do this in <a href="http://www.bioclipse.net/">Bioclipse2</a> too?</li>
  <li>Can we improve the default CDK descriptor parameters to maximize the column count?</li>
  <li>Rajarshi, what would be involved to write some wrapper code for atomic descriptors for rcdk?</li>
</ul>]]></content><author><name>Egon Willighagen</name></author><category term="cdk" /><category term="defense" /><category term="phd" /><category term="metabolomics" /><category term="cheminf" /><category term="chemometrics" /><category term="justdoi:10.18637/jss.v018.i02" /><summary type="html"><![CDATA[As announced earlier , Miguel, Velitchka, Christoph and I held a small CDK/Metabolomics/Chemometrics unconference. We started late, and did not have an evening program, resulting in not overly much results. However, we did do molecular chemometrics.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/riPred.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/riPred.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">T plus 18 hours: dr and preparing for the afterparty, umm ^w^w^w, CDK/Metabolomics/Chemometrics unconference</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/04/03/t-plus-18-hours-dr-and-preparing-for.html" rel="alternate" type="text/html" title="T plus 18 hours: dr and preparing for the afterparty, umm ^w^w^w, CDK/Metabolomics/Chemometrics unconference" /><published>2008-04-03T00:00:00+00:00</published><updated>2008-04-03T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/04/03/t-plus-18-hours-dr-and-preparing-for</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/04/03/t-plus-18-hours-dr-and-preparing-for.html"><![CDATA[<p>I am doctor now; I shall now be <a href="http://taaladvies.net/taal/advies/tekst/21#6">addressed as</a> <em>weledelzeergeleerde</em> Egon;
translating to something like <em>quite-noble-very-knowledgeable</em>, hahahaha. I’ll put up a few photo’s of the ceremony, which
is actually quite formal at the <a href="http://www.ru.nl/">Radboud University</a>, later.</p>

<p>With this blog item, I would to thank everyone who left a message, sent email, etc with good luck messages. Very much
appreciated! I’d also like to thank my supervisors, promotores <a href="http://www.cac.science.ru.nl/people/lbuydens/index.html">Lutgarde Buydens</a> and
<a href="http://wwmm.ch.cam.ac.uk/blogs/murrayrust/">Peter Murray-Rust</a> (he mentions the event <a href="http://wwmm.ch.cam.ac.uk/blogs/murrayrust/?p=1019">here</a>),
and <a href="http://www.cac.science.ru.nl/people/rwehrens/index.html">Ron Wehrens</a> for their confidence in me and their guidance
on the path towards the post-doc life. I also thank all those who attended my defense; I had a brilliant day, and actually
enjoyed talking to those who took place in my promotion committee and who asked me the not-really-nasty-questions about
my work.</p>

<h2 id="cdk-chemometrics-in-metabolomics-unconference">CDK-Chemometrics in Metabolomics Unconference</h2>

<p>For today, I organized a small, informal <a href="http://en.wikipedia.org/wiki/Unconference">unconference</a>, oriented around the
<a href="http://cdk.sf.net/">CDK</a>, chemometrics and metabolomics. I’m certain we will be online much of the day, as we typically
do. The meeting will start around 10:00 <a href="http://en.wikipedia.org/wiki/Central_European_Summer_Time">CEST</a>, but we’ll
attend a seminar by <a href="http://www.ki.si/index.php?id=844">Marjana Novič</a> at 11:00 CEST. If you happen to be in
<a href="http://en.wikipedia.org/wiki/Nijmegen">Nijmegen</a>, just drop in on the Analytical Chemistry department.
Otherwise, join the #cdk chat channel in the irc.freenode.net network.</p>

<p>What we’ll do?? Hey, it’s an unconference; we have no idea yet :)</p>]]></content><author><name>Egon Willighagen</name></author><category term="defense" /><category term="cheminf" /><category term="chemometrics" /><category term="phd" /><summary type="html"><![CDATA[I am doctor now; I shall now be addressed as weledelzeergeleerde Egon; translating to something like quite-noble-very-knowledgeable, hahahaha. I’ll put up a few photo’s of the ceremony, which is actually quite formal at the Radboud University, later.]]></summary></entry><entry><title type="html">T minus 26 hours: defending open source chemoinformatics (and more)</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/04/01/t-minus-26-hours-defending-open-source.html" rel="alternate" type="text/html" title="T minus 26 hours: defending open source chemoinformatics (and more)" /><published>2008-04-01T00:00:00+00:00</published><updated>2008-04-01T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/04/01/t-minus-26-hours-defending-open-source</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/04/01/t-minus-26-hours-defending-open-source.html"><![CDATA[<p>In about 26 hours from now, I will be <a href="https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work.html">defending my PhD thesis <i class="fa-solid fa-recycle fa-xs"></i></a>.
Follow that link to read the summary; I was thinking if publishing my introduction and discussion (the rest has been published in peer-reviewed
journals) on <a href="http://precedings.nature.com/">Nature Precedings</a>; would that be a good idea? Otherwise, I’ll post it in my blog. If you just
happen to want to attend the public defense, it’s
<a href="http://maps.google.com/maps?f=q&amp;hl=en&amp;geocode=&amp;q=Comeniuslaan+2,+6525+Nijmegen,+Nijmegen+(Gelderland),+Netherlands&amp;sll=37.0625,-95.677068&amp;sspn=28.114729,75.234375&amp;ie=UTF8&amp;ll=51.820699,5.857548&amp;spn=0.002673,0.009184&amp;t=h&amp;z=17&amp;iwloc=addr">here</a>:</p>

<p><img src="/assets/images/aula.png" alt="" /></p>]]></content><author><name>Egon Willighagen</name></author><category term="cheminf" /><category term="chemometrics" /><category term="phd" /><summary type="html"><![CDATA[In about 26 hours from now, I will be defending my PhD thesis . Follow that link to read the summary; I was thinking if publishing my introduction and discussion (the rest has been published in peer-reviewed journals) on Nature Precedings; would that be a good idea? Otherwise, I’ll post it in my blog. If you just happen to want to attend the public defense, it’s here:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/aula.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/aula.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">TODO: April 2nd, defend my PhD work</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work.html" rel="alternate" type="text/html" title="TODO: April 2nd, defend my PhD work" /><published>2008-03-01T00:00:00+00:00</published><updated>2008-03-01T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work.html"><![CDATA[<p>In 4.5 weeks, on Wednesday April 2 (13:30 precisely, <a href="http://maps.google.com/maps?f=q&amp;hl=en&amp;geocode=&amp;q=comeniuslaan+2,+nijmegen,+nederland&amp;sll=37.0625,-95.677068&amp;sspn=25.010803,75.234375&amp;ie=UTF8&amp;ll=51.820852,5.857548&amp;spn=0.002374,0.009184&amp;t=h&amp;z=17&amp;iwloc=addr">Aula, Comeniuslaan 2, Nijmegen</a>)
I will publicly defend my PhD work performed in the <a href="http://www.cac.science.ru.nl/">Analytical Chemistry group</a> of
<a href="http://scholar.google.nl/scholar?as_q=&amp;num=10&amp;btnG=Search+Scholar&amp;as_epq=&amp;as_oq=&amp;as_eq=&amp;as_occt=any&amp;as_sauthors=LMC+Buydens&amp;as_publication=&amp;as_ylo=&amp;as_yhi=&amp;as_allsubj=all&amp;hl=en&amp;lr=">Prof. Lutgarde Buydens</a>
at the <a href="http://www.ru.nl/">Radboud University Nijmegen</a>:</p>

<p><img src="/assets/images/thesisCover.png" alt="" /></p>

<h2 id="table-of-contents">Table of Contents</h2>

<ol>
  <li>Introduction</li>
  <li>Molecular Chemometrics (doi:<a href="https://doi.org/10.1080/10408340600969601">10.1080/10408340600969601</a>)</li>
  <li>1D NMR in QSPR(doi:<a href="https://doi.org/10.1021/ci050282s">10.1021/ci050282s</a>)</li>
  <li>Comparing Crystals (doi:<a href="https://doi.org/10.1107/S0108768104028344">10.1107/S0108768104028344</a>)</li>
  <li>Supervised SOMs (doi:<a href="https://doi.org/10.1021/cg060872y">10.1021/cg060872y</a>)</li>
  <li>Chemical Metadata in RSS (doi:<a href="https://doi.org/10.1021/ci034244p">10.1021/ci034244p</a>)</li>
  <li>Interoperability (doi:<a href="https://doi.org/10.1021/ci050400b">10.1021/ci050400b</a>, the Blue Obelisk paper)</li>
  <li>Discussion and Outlook</li>
</ol>

<p>Chapters 2, 3, 4, and 5 are first author papers, while for chapters 6 and 7 I am just co-author.</p>

<h2 id="summary">Summary</h2>

<p>Chemometrics and chemoinformatics play important roles in the analysis and modeling of molecular data. In particular, in understanding and
prediction of properties of molecules and molecular systems. Both chemometrics and chemoinformatics apply statistics, machine learning and
informatics methodologies to chemical questions, though originating from a different background. Where chemometrics had its origins in the
extraction of information from chemical experiments, chemoinformatics had roots in the representation of chemical data for storage in
databases. The technological advances in chemistry and biochemistry in the past decades have led, however, to a flood of data and new
questions, and the data analysis and modeling have become more complex. The standing challenge in data analysis and data exchange, is how
to represent the molecular features relevant to the problem at hand. This representation of molecular information is the topic of this
thesis.</p>

<p>Chapter 1 introduces the field of data analysis and modeling of molecular data and describes the aforementioned importance of representation
of relevant features. It discusses different approaches to molecular representation, such as line notations, chemical graphs, and quantum
chemical models. Each of these have limitations when used in data analysis and modeling. Numerical representations are then introduced, which
allow the application of statistical and mathematical modeling approaches. These numerical representations are commonly derived from chemical
graph and quantum chemical representations. CoMFA and the classification of enzyme reactions are examples were the choice of molecular
representation as well as the analysis method are important.</p>

<p>The term <em>molecular chemometrics</em> is coined in Chapter 2 for the field that applies statistical modeling methods to molecular structure.
It reviews the advances made in this field in recent years. New numerical descriptors for molecules are discussed, as well as approaches to
represent molecules in more complex systems like crystal structures and reactions. Molecular descriptors are used in similarity and diversity
analysis. The applications of new methods for structure-activity and structure-property modeling, and dimension reduction are described. An
overview of recent approaches in model validation show new insights and approaches to estimate the performance of classification and regression
models. The last section of this chapter lists new databases and introduces new methods that improve the extracting of chemical data from
database and repositories. Semantic markup languages improve the exchange of data, and new methods have been introduced to extract molecular
properties from text documents.</p>

<p>Chapter 3 studies the in literature proposed use of 1D <sup>13</sup>C and <sup>1</sup>H NMR spectra as molecular descriptor. These spectra
are known to describe features relevant to physical properties like solubility and boiling point. The NMR representation is studied for the
predictive powers of its PLS models for three structure-property data sets. The results indicate that proton NMR is not suitable for building
QSPR models in combination with PLS. Carbon NMR-based models, however, do give reasonable QSPR models, and the regression vectors for the
carbon NMR data, correlate with spectral regions relevant to molecular fragments. Nevertheless, the predictive power of the carbon NMR-based
spectra is still less than models based on common molecular descriptors. It is concluded that NMR spectra should not be considered first
choice when making predictive models in general, and that proton NMR should probably not be used at all.</p>

<p>A computational method to calculate similarities between crystal structures based on a new representation is introduced in Chapter 4. While
a reference method is perfectly able to identify structures with high similarity, it fails to recognize the different similarities between
two similar structures and two completely different structures. This makes it very difficult for clustering algorithm to organize small
clusters of identical and highly similar structures into larger clusters. The new representation of crystal structures introduced in this
chapter shows a much smoother transition in similarity values when crystal structures go from identical, via similar, and finally to
dissimilar structures. Clustering a set of simulated polymorphic structures of estrone, and classification of a set of experimental
cephalosporin structures reproduce expected clustering and classification.</p>

<p>Chapter 5 uses supervised self-organizing maps to cluster crystal structures represented by their powder diffraction pattern and one or
more properties. The topological structure of the resulting maps not only depends on the similarity of the diffraction data, but also on
the properties of interest, such as cell volume, space group, and lattice energy. This approach is used to analyze and visualize large
sets of crystal structures, and the results show that these supervised maps not only give a better mapping, they can also be used to predict
crystal properties based on the diffraction patterns, and for subset selection in polymorph prediction. The two applications in
crystallography show that suitable representations and similarity measures that allow data analysis and modeling of molecular crystal data
are now available. Both approaches are flexible enough to open up a new field of research; especially combinations with other classification
schemes for crystal structures, such as those based on hydrogen bonding patterns, come to mind.</p>

<p>Chapter 6 introduces and discusses a method that allows information rich distribution of molecular data between machines, such as measuring
devices and computers. Existing approaches often imply not or badly documented semantics which may lead to information loss. CMLRSS is
proposed and combines two existing web standards: Rich Site Summaries (RSS), also known as RDF Site Summaries, and the Chemical Markup
Language (CML). Here, RSS is used as transport layer, while CML is used to contain the chemical information. CML supports a wide range of
chemical data, including molecular (crystal) structures, reaction schemes, and experimental data such as NMR spectra. It is shown that
this semantic representation allows automated dissemination of chemical data, and is increasingly used to exchange data between web
resources.</p>

<p>Chapter 7 describes a communal effort to realize interoperability in chemical informatics, which is called the Blue Obelisk movement.
This movement currently consists of more than ten smaller and larger, open source and open data projects all related to chemoinformatics
and chemistry in general. To increase the reproducibility of molecular representations, this chapter introduces a collaborative dictionary
of chemoinformatics algorithms, and a public repository of chemical data of general interest, including data for chemical elements and
isotopes, (boiling points, colors, electron affinities, masses, covalent radii, etc.), definitions of atom types, and more. The
availability of a standard set of atomic properties, open source algorithms and open data (for example via CMLRSS feeds), it is much
easier to reproduce and validate published results in molecular chemometrics. Results from Chapter 3 show that such ability is no luxury.</p>

<p>The last chapter summarizes the efforts in this thesis and how they address the challenges in molecular chemometrics. This thesis shows
the strong interaction between representation and the methods used for data analysis: molecular representation need to capture relevant
information and be compatible with the statistical methods used to analyze the data. The chapters review molecular
representations and put focus on model validation using statistics, visualization methods, and standardization approaches.</p>]]></content><author><name>Egon Willighagen</name></author><category term="cheminf" /><category term="chemometrics" /><category term="phd" /><category term="doi:10.1080/10408340600969601" /><category term="doi:10.1021/CI050282S" /><category term="doi:10.1107/S0108768104028344" /><category term="doi:10.1021/CG060872Y" /><category term="doi:10.1021/CI034244P" /><category term="doi:10.1021/CI050400B" /><summary type="html"><![CDATA[In 4.5 weeks, on Wednesday April 2 (13:30 precisely, Aula, Comeniuslaan 2, Nijmegen) I will publicly defend my PhD work performed in the Analytical Chemistry group of Prof. Lutgarde Buydens at the Radboud University Nijmegen:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/thesisCover.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/thesisCover.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>