Plate VIMethodology
Open Science, Rigorous Method
How the platform is built — the standards we build on, the data-quality protocols we follow, and the open-science commitments that make the research reproducible and citable.
⟢ Open Scholarly Infrastructure
FAIR Data and POSI Commitments
The platform follows the FAIR Principles (Findable, Accessible, Interoperable, Reusable) and aligns with POSI — the Principles of Open Scholarly Infrastructure — to ensure the data layer does not become dependent on any single institution or commercial provider.
- Citable dataset releases with DOIs through Zenodo-class archiving — each release is permanently citable and versioned.
- ORCID integration for contributor identity — every individual contribution is attributed, not just institutional credit.
- No vendor lock-in for the data infrastructure — stack choices follow open standards with documented migration paths.
- The bibliographic-search protocol is documented and reproducible — literature reviews can be independently replicated.
⟢ Interoperability Standards
Standards We Build On
Every architectural decision in the data layer traces back to an established, community-governed standard. This is not incidental — it is the mechanism by which the platform remains interoperable with external infrastructure and trustworthy over decades.
WoRMS / AphiaID
World Register of Marine Species provides the authoritative taxonomic spine. Every species record is keyed by AphiaID — a permanent identifier that enables WoRMS synchronisation and prevents taxonomic bifurcation.
Darwin Core
The TDWG exchange standard governs how occurrence and taxon data are packaged for interoperability with GBIF, OBIS, and institutional repositories.
ColDP
Catalogue of Life Data Package defines the exchange format for taxonomic checklists. It aligns our internal taxon model with international aggregators.
SKOS
Simple Knowledge Organisation System structures the controlled morphological vocabulary: prefLabel canonicalises a term; altLabel captures variant spellings and language translations across the literature.
pgvector + OpenSearch
Vector embeddings (pgvector on Supabase) support semantic similarity retrieval; OpenSearch drives the faceted morphological filter layer of the Testaria System.
FAIR Principles
Every dataset is designed to be Findable, Accessible, Interoperable, and Reusable. Citable releases carry DOIs (Zenodo-class archiving) and contributor identity is tracked via ORCID.
⟢ Scientific QC Protocol
Data Quality & Reliability
Scientific credibility depends on explicit quality commitments, not just good intentions. Four non-negotiable principles govern how every datum enters and lives in the catalog.
- Provenance
- Every datum carries a reference to its source publication. Nothing enters the catalog without an auditable bibliographic anchor.
- Morphology ≠ interpretation
- Observable features (chamber count, coiling direction, aperture form) are recorded separately from taxonomic conclusions. The two layers never conflate.
- Controlled vocabulary
- A SKOS vocabulary reconciles lexical variation across the literature — "foliate" and "foliáceo" resolve to the same prefLabel, preserving traceability.
- Inter-rater agreement
- A scientific QC protocol (inter-rater agreement study) underpins reliability scoring per identification. Disagreements are flagged, not silently averaged.
Every identification record carries a reliability score derived from inter-rater agreement — disagreements are flagged explicitly rather than silently averaged.
⟢ Source Reconciliation
Three-Source Reconciliation
No single source covers the full taxonomic record. The catalog is built by reconciling three complementary sources; each plays a defined, non-overlapping role.
Licensing note — Ellis & Messina Catalogue
The Ellis & Messina Catalogue of Foraminifera is a licensed resource. Original taxonomic descriptions from the catalogue are used strictly for internal enrichment — to inform vocabulary alignment and QC — and are never redistributed through the platform. Only our derived, structured metadata is published.
WoRMS
Taxonomic spine
Authoritative, continuously maintained. AphiaID is the canonical key — the record is always synced, never forked.
Ellis & Messina Catalogue
Legacy morphological index
Internal enrichment only. Original descriptions are licensed and are never redistributed. The catalogue informs vocabulary alignment, not public data.
Primary literature
Morphological observations
Systematically ingested via the Nomia pipeline. Every datum carries a bibliographic provenance record pointing back to the source publication.
⟢ Literature Pipeline
Systematic Bibliographic Search
The Nomia Library is fed by a systematic, reproducible literature-search protocol. Search strings, source databases, inclusion/exclusion criteria, and retrieval dates are documented so that any researcher can independently replicate or update the search. Retrieved publications flow through Nomia's extraction pipeline — PDF-to-structured-data — and candidate morphological observations undergo human review before entering the catalog.
Indexing sources include major scholarly databases (OpenAlex, CORE, Semantic Scholar, OpenCitations) and specialist repositories. Citation backtracking and forward citation tracking are used to recover works missed by keyword search alone.