Plate VDevelopment Plan
A phased path from taxonomic spine to global infrastructure.
The platform follows a Foundation → Core → Validation → Scale → Ecosystem lifecycle. Phases are not strictly sequential — Nomia and Testaria develop in parallel. Current status: Pre-MVP.
- 1–2Phase 1–2In progress
Taxonomic spine and data layer
Establish the foundational standards: AphiaID-keyed WoRMS synchronisation, the character vocabulary, and a pilot dataset of known species. Prove the concept end-to-end on a bounded cohort before scaling.
- WoRMS AphiaID sync — forams_* → foram2_* → OpenSearch (ingested)
- WoRMS MCP server built and operational
- Character vocabulary v0 (P0.2) — seeded from Foraminifera.eu schema + Loeblich & Tappan
- Data model v0.1 (P0.1) — entities, provenance, morphology / interpretation boundary
- Pilot cohort V1 — shallow-platform benthics, extant material
- Synonymy resolution via gnparser / gnverifier + ColDP model
- 3Phase 3In progress
Morphological search on real data
Take the Testaria identification filter from a mock-data demo to a working system backed by the real character matrix. A minimum viable character set (~5–8 characters) with controlled-state vocabulary is enough to prove reproducible identification.
- Testaria MVP demo running on mock data (in progress)
- Minimum character set and controlled state vocabulary (in design)
- Full Supabase + OpenSearch query layer (in design)
- Faceted filter UI and ranked candidate output
- CSV export — species list + references + reliability record
- Methodological article + pilot dataset DOI (Zenodo)
- 4Phase 4Planned
Literature → structured data pipeline
Industrialise the bibliographic ingestion pipeline: acquire publications, convert PDFs to structured records, extract morphological character candidates via NLP, and populate the reference and species-description banks.
- Bibliographic acquisition — OpenAlex / CORE / Semantic Scholar / Zotero
- PDF → Markdown conversion (GROBID / Docling / Nougat)
- Species-name detection (GNfinder) and metadata extraction
- Character candidate extraction with LLM assist + human review
- Community contribution protocol and co-curation staging queue
- 5Phase 5Planned
Ecology layer and open data exports
Add the ecology and distribution dimension, and open the platform fully to the global biodiversity network. Citable releases make the catalogue a referenceable scientific object.
- EcologyDistribution entity — occurrence / environmental parameters
- Darwin Core / ColDP export for GBIF and OBIS
- BENFEP, ForCenS, and BFR2 dataset integration
- Public REST / GraphQL / MCP API for human and agent consumption
- Webhooks for partner ingest and release notification
- Citable versioned releases with DOI (Zenodo)
- 6Phase 6Planned
Education and communication layer
Build the public-facing educational layer that converts the dense scientific catalogue into structured learning materials. Foram World is conditioned on data maturity — it only becomes meaningful once the catalogue has sufficient depth.
- Species educational cards — visual guides, morphological diagrams
- Blog and curated publication feed (Sanity-managed content)
- Highlighted terms and recently indexed taxa
- Structured learning pathways for new researchers
- Direct links to species records and identification tool
- 7Phase 7Planned
High-quality image layer and 3D visualisation
Expand the media layer with curated photomicrographic archives and optional 3D specimen visualisation. Image-based identification (CNN) may enter as a complementary future layer — not the core approach.
- Specimen image ingest with per-image licence governance
- Visual comparison panel alongside the morphological filter output
- 3D specimen visualisation (open decision)
- Image-based classification layer (future, not core)
⟢ Scientific Publications
Publications timed to build credibility and attract funding.
Publications are a strategic instrument, not an afterthought. Each phase produces a citable scientific record that registers progress and opens doors to institutional funding.
Phase 1–2
Morphological filter methodology
Methodological article + pilot seed dataset (DOI via Zenodo) + congress abstract
Phase 3–4
Nomia pipeline + ecology integration
Pipeline paper + Darwin Core / FAIR integration description
Validation (Q1/Q2)
Inter-rater concordance study
Explicit reliability methodology and confidence-scoring validation
Continuous
Citable catalogue releases
Versioned DOI releases as the catalogue grows; conference dissemination once minimum species count is reached
5–10+ year horizon · Open Science
The platform does not end — it updates perpetually.
Each phase stabilises the standards that unlock the next. Once the pilot cohort forges the vocabulary, all subsequent taxonomic classes can scale against it. The mission is long; the path is clear.