The dbDNA project aims at providing a system for reliable taxonomic annotation of DNA barcoding and metabarcoding sequences by (1) rating reference sequences based on various criteria, (2) employing expert identifier verification of specimens and (3) filtering out sequences that are erroneous. Additionally, (4) cases where unambiguous identification through (meta)barcoding fails for reasons such as hybridisation and incomplete lineage sorting are collected.
(1) By employing the newly developed dbSCORE Pipeline ( https://github.com/TillMacher/dbDNA), reference sequences for any given taxa list can be downloaded from the Barcode of Life Data System database ( https://boldsystems.org/) and all individual sequences will be graded based on different criteria such as expert identifier verification (+15 points), taxonomic concordance (+15), geography (+2-4), life stage (+1) and sex (+1) of reference specimens (see publication later this year). Ultimately, sequences will be categorised as gold (≥40 points), silver (25-39), bronze (10-24), and unreliable (≤9) depending on their total point value. Please note that no sequence can be of “gold standard” without the expert identifier verification and taxonomy concordance criterion.
(2) The points for “expert identifier verification” will be granted if the respective identifier of a reference specimen is listed on our identifier expert list,
which follows a white list approach and can be found here: The expert identifier list is not yet available.
In case you want to be listed on the whitelist, please contact Arne Beermann (arne.beermann@uni-due.de).
(3) Points for “taxonomic concordance” will be granted if only one species name is associated with a cluster of sequences (here based on genetic distance).
This is often not the case in public databases, for instance due to misidentified reference specimens or usage of synonyms. In obvious cases, these “discordant sequences” will be blocked from the rating process, ultimately improving the reliability of taxonomic annotation.
Please note, that not all blocked sequences and reference specimens have to be ultimately false. All currently blocked sequences can be found on the block list: dbDNA_Blocklist_260603_FWE.xlsx (88.2 KB)
In case you have a comment on one or more blocked sequences, or know of additional cases that should be blocked, please contact Arne Beermann (arne.beermann@uni-due.de).
(4) Lastly, there are cases such as hybridisation or incomplete lineage sorting where DNA barcoding and metabarcoding will fail in identifying species. However, as a community we can make an effort of collecting such cases for correct interpretation of sequencing results.
The issue list is a starting point to collect such cases: dbDNA_Issuelist_260625_FWE.xlsx (17.0 KB)
In case you know of additional cases to complement the list (please give a reference if possible), please contact Arne Beermann (arne.beermann@uni-due.de).
Filling the expert identifier permit list and issue list will only fill over time and heavily benefit from activity of and engagement with taxonomic experts and the (meta)barcoding community.
Both lists will be updated regularly and available as versioned files.
Complementing and curating the sequence block list is the main curation step of the proposed system. Polyphyletic sequence clusters have to be checked manually and individual erroneous sequences have to be blocked accordingly.
This can and will happen only over time by taking different taxa lists as input for the dbSCORE pipeline. As a first taxa list to provide and curate reliable reference sequences,
we used the Operational Taxa List (OTL) for assessing the ecological status of German water bodies: Operational Taxa List, Germany (OTL)