ChemSpider data cleanup

15 Dec 2023

In previous posts, we have discussed the automated workflow we use to check new incoming data for structure and synonym errors. These checks allow us to remove the most common types of errors before they are added to the site. However, these filters do not apply to data already in ChemSpider.

Manual curation is an important part of our work. We periodically review the data on our most accessed records, in addition to ad-hoc removal or correction of erroneous data that we or our users notice when using the site. However, there are far too many records and far too much data to clean up using manual curation alone.

Recently we have focused on bulk identification and removal of erroneous data. This work has covered mapping errors and other clearly incorrect values in our experimental property data, correction or removal of malformed synonyms, correction of incorrectly labelled synonyms, and resolution of structure/synonym clashes.

Experimental Properties

We retrieved all 6.3 million experimental properties, text properties, and associated annotations from the ChemSpider database. We then compared the original text of the property as it was written in the original file to how that text was parsed and mapped by our deposition system. This enabled us to identify and correct several types of errors affecting around 2% of the properties in our database:

35,774 experimental property values had been assigned the incorrect unit (e.g. g/L instead of g/mL, °C instead of °F)
2,591 boiling points measured under non-standard pressure did not have this pressure displayed
4,292 densities had their density and temperature values swapped
79,252 miscellaneous erroneous properties and associated annotations were deleted. For example, “white crystals” mapped as melting point, impossibly high melting points or densities, etc.

Synonyms

Synonyms, chemical names, and identifiers are the most abundant type of data on ChemSpider, with a total of more than 446 million synonyms. These synonyms have additional metadata including language labels and flags identifying what type of synonym they are (e.g. CAS number, UNII, INN, trade name).

Simple Checks

We ran a series of regular expression string searches to identify synonyms with incorrect metadata, as well as malformed or otherwise erroneous synonyms.

200,007 synonym type flags added, and 4,766 incorrect flags removed
9,170 synonyms with an incorrect language label identified.
631,697 erroneous synonyms identified, including scrambled characters, properties/units, molecular formulae as synonyms, purity information, or invalid CAS numbers or EC numbers (formerly called EINECS).
922,334 instances of these erroneous synonyms deleted from ChemSpider records.

Structure/Synonym comparison

After identifying and removing these synonym-level errors, we then cross-checked ChemSpider records and their synonyms to identify mismatches. This work included amino acids, nucleic acids, and pharmaceutically acceptable salts.

As a first pass, we compared synonyms to molecular formulae to identify records missing key elements. Examples include synonyms describing a sodium salt when the molecular formula does not contain sodium, or describing an amino acid when the molecular formula contains no nitrogen. A total of 28,194 of these synonym/formula clashes were identified and removed.

For records that passed this initial molecular formula check, we performed a SMARTS comparison to identify chemical structures missing key structural features described in the synonym. These SMARTS strings were written broadly, with common substitutions allowed to prevent unnecessary removal of valid synonyms from derivative compounds.

In the following examples, the mismatched part of the synonym is highlighted in bold.

Structure	Removed synonym
	Sulfate ion
	Zolpidem tartrate
	Sodium S-sulfocysteine hydrate

After identifying these clashes, we manually spot-checked the output to weed out false positives and iterate the SMARTS filters. 101,257 synonym/structure clashes were identified and removed.

These checks included the following categories:

Amino acids and their derivatives: 6 formula clashes, 56 structure clashes
Nucleic acids, nucleosides, nucleotides: 977 formula clashes, 1,870 structure clashes
Halogens: 13,437 formula clashes, 1,256 structure clashes
Alkali and alkaline earth metals, and aluminium: 3,586 formula clashes, 56 structure clashes
Carboxylic acids and their derivatives: 5,002 formula clashes, 88,501 structure clashes
Other pharmaceutically acceptable acids: 3,534 formula clashes, 1,529 structure clashes
Amides and amines: 190 formula clashes, 304 structure clashes
Deuterates, hydrates, methylbromides: 1,462 formula clashes, 7,685 structure clashes

Get involved

You are the expert in your area of chemistry, so if you see something that doesn’t look quite right please let us know. If the error is confined to a single ChemSpider record, click the “Comment On This Record” box at the top of the affected record and let us know what the problem is. All we need is a sentence describing the error, however the more information you can provide, the better.

For more systemic errors, or in cases where you want to attach supplementary information or corrected chemical structures, please get in touch via email (chemspider@rsc.org).

Comments Off on ChemSpider data cleanup

Webinar 3: Chemistry data: Challenges and opportunities. Watch the recording

18 Nov 2023

By Richard Kidd.

We will explore ongoing and planned initiatives developing standards and tools, research infrastructures, and cultures to support FAIR chemistry data as well as its preparation, publication, and reuse.

Webinar 3: Challenges and opportunities

Webinar recorded on 7 December 2023 – watch the recording here

Speakers

“How to initiate the cultural change towards digital chemistry” SLIDES
Sonja Herres-Pawlis
Chair of Bioinorganic Chemistry, RWTH Aachen

“How can we combat heterogeneous, unfair and disparate data in digital chemistry? ” SLIDES
Samantha Kanza
Senior Enterprise Fellow, University of Southampton
Pathfinder Lead, Physical Sciences Data Infrastructure (PSDI)

“How data journals can support (chemistry) data sharing and discovery” SLIDES
Guy Jones
Chief Editor of Scientific Data, Springer Nature

Sponsored by Revvity

Revvity Signals Software, formerly PerkinElmer Informatics, has over three decades of experience providing support for scientific workflows.

Our powerful informatics solutions are used in R&D across disciplines from drug discovery to materials development. Now under our Signals Research Suite, our end-to-end SaaS solution integrates workflows to accelerate innovation and help scientists collaborate. In addition, our solution powered by TIBCO® Spotfire® can transform clinical trials.

From our flagship ChemDraw® and E-Notebook applications, to our Signals Research Suite, to our TIBCO® Spotfire® partnership for data analytics, Revvity Signals offers a powerful suite of scientific solutions.

Supported by

About ChemSpider

Explore more than 128 million structures on the ChemSpider database. Including over 200 data sources, ChemSpider is a valuable source of information for chemical scientists working with data.

Freely accessible and comprehensive, this rich source of structure-based chemistry information is a fundamental resource for chemical scientists working with data everywhere.

Learn more about ChemSpider

Comments Off on Webinar 3: Chemistry data: Challenges and opportunities. Watch the recording

Webinar 2: What does the future hold? Watch the recording

17 Oct 2023

By Richard Kidd.