Databases queried
29 sources are registered. Those marked needs key stay
dormant until you add a credential in config.php; the scan runs without them, just
with less curated evidence.
Tier 1 — Culture collections and nomenclatural authorities weight 5
Curated by the collections that actually hold the strains. Weighted highest — a medium recorded here was used to keep the organism alive.
| Source | Access | Status | What it contributes |
|---|---|---|---|
| DSMZ MediaDive ↗ | API | active | Authoritative growth-media recipes and strain-to-medium links. |
| BacDive (DSMZ) ↗ | API | needs key | Curated strain phenotypes: medium, temperature, Gram, morphology, enzymes. |
| LPSN (DSMZ) ↗ | API | needs key | Nomenclature, type strains, effective/valid publication. |
| StrainInfo ↗ | API | active | Cross-collection strain equivalences and deposit history. |
| NCBI Taxonomy ↗ | API | active | Accepted lineage; supplies the phylum prior for Gram inference. |
| ATCC Catalogue ↗ | deep link | manual | Product sheets carry medium and temperature. No open API; deep link only. |
| NCTC / UKHSA Culture Collections ↗ | deep link | manual | Reference strain sheets with recommended media. Deep link only. |
Tier 2 — Peer-reviewed indexed literature weight 3
Published, refereed work. The bulk of the evidence.
| Source | Access | Status | What it contributes |
|---|---|---|---|
| PubMed (NCBI E-utilities) ↗ | API | active | Titles and abstracts; the primary evidence stream. |
| PubMed Central (open access) ↗ | API | active | Open-access full text — the richest source of Methods-section detail. |
| Europe PMC ↗ | API | active | Abstracts plus OA full text, including non-PubMed content. |
| Europe PMC full text ↗ | API | active | Fetches Methods sections for OA hits found in the step above. |
| Crossref ↗ | API | active | DOI metadata and abstracts where publishers deposit them. |
| OpenAlex ↗ | API | active | Open scholarly graph; inverted abstracts are reconstructed locally. |
| Semantic Scholar ↗ | API | active | Abstracts and citation context. Keyless access is rate-limited. |
| DOAJ ↗ | API | active | Open-access journal articles, strong on regional microbiology titles. |
| OpenAIRE Explore ↗ | API | active | European research outputs aggregated across repositories. |
| SciELO ↗ | API | active | Latin-American and Iberian journals under-represented elsewhere. |
| CORE ↗ | API | needs key | Aggregated repository full text. |
| Springer Nature Meta API ↗ | API | needs key | Includes Bergey-adjacent monograph and journal metadata. |
| Elsevier ScienceDirect ↗ | API | needs key | Institutional key required; respects entitlement. |
| NCBI Bookshelf ↗ | API | active | Medical Microbiology (Baron), StatPearls and similar reference texts. |
| Google Scholar ↗ | deep link | manual | No public API and automated querying is against its terms. The app builds a search link for you to open manually — it does not scrape it. |
Tier 3 — Preprints, repositories and protocols weight 1.5
Useful and often current, but unrefereed. Weighted down so it can support a value but rarely decide one.
| Source | Access | Status | What it contributes |
|---|---|---|---|
| bioRxiv / medRxiv (via Europe PMC) ↗ | API | active | Preprint server content. Not peer reviewed — weighted down accordingly. |
| Zenodo ↗ | API | active | Datasets, theses and protocols deposited openly. |
| protocols.io ↗ | deep link | manual | Step-by-step culture protocols. Deep link only. |
| ResearchGate ↗ | deep link | manual | Crawling disallowed. Deep link only. |
Tier 4 — Encyclopedic reference weight 1
Orientation only. Never decides a field on its own.
| Source | Access | Status | What it contributes |
|---|---|---|---|
| Wikipedia ↗ | API | active | Orientation only. Lowest weight; never decides a field on its own. |
| Wikidata ↗ | API | active | Structured identifiers used to bridge to other catalogues. |
| GBIF ↗ | API | active | Name resolution and synonymy cross-check. |
On the two we do not crawl
Google Scholar has no public API and its terms prohibit automated querying; ResearchGate likewise disallows crawling. Both are registered as deep links — the app composes the query string and hands you the URL. It never fetches them. Any tool that claims to scrape Scholar at scale is either using an undocumented proxy that breaks weekly, or getting the host IP blocked.
The gap that leaves is real, and the honest way to close it is credentials: BacDive and LPSN both give free accounts, and BacDive in particular is curated exactly for these seven fields.