About GanoDB
GanoDB is a genome, annotation and applied-research database for Ganoderma boninense, the basidiomycete that causes basal stem rot of oil palm. It is built on the chromosome-level reference assembly GCA_002900995.3 (Gabo_G3, strain G3) — 55.87 Mb across 12 pseudo-chromosomes.
What it holds
- 16,748 protein-coding genes and 3,573 non-coding RNA genes, at 96.3% BUSCO (basidiomycota_odb10).
- 16,141 ESMFold structures (96.4% of the proteome), all Amber-minimised.
- Expression from 39 quantified public RNA-seq runs across four condition groups.
- Core/accessory content across three admissible isolates — 13,003 core orthogroups, 3,487 accessory, 3,367 strain-unique.
- Comparative placement against eight species, and 1.6 M variants across four G. boninense isolates.
- Transferred evidence: 4,927 PHI-base phenotypes over 4,428 proteins, 80,315 STRING-derived interactions, and 147 species papers.
How the annotation was built
Gene models come from BRAKER3 (RNA-seq + protein evidence) with balanced-TSEBRA recovery and miniprot rescue. Function is assigned by Pfam/InterPro, dbCAN (CAZymes), MEROPS (proteases), eggNOG (COG, GO, KEGG KO), EC, TCDB (transporters), SignalP 6 and TMHMM (topology), EffectorP 3 (candidate effectors), antiSMASH (biosynthetic clusters) and DIAMOND against NCBI NR. Sequence-dark proteins are probed structurally — ESMFold model, then Foldseek and DeepFRI — and those calls are labelled hypothesis-grade wherever they appear.
Known limits
- 607 of the 16,748 protein-coding genes have no predicted structure. Their proteins exceed what the available GPUs can fold in one pass — the longest is 5,047 aa.
- 5,228 orthogroups carry an absence that cannot be checked — they contain no G3 protein to use as a tblastn query. Since ~22–27% of checkable absences were refuted at genome level, the true core is a lower bound and the unique compartment an upper one.
- Two of the five public assemblies are not usable for absence. NJ3 is fragmented (N50 6.1 kb) and T10 is an uncollapsed dikaryon, so for those a missing gene is a statement about the assembly.
- Six of the 45 public RNA-seq runs carry no TPM. They feed gene prediction and transcript assembly but are not in the expression matrix; they are badged as such on the sequencing page.
- A previously published core figure of 13,808 orthogroups was withdrawn. A tblastn hit falling inside an already-annotated gene is a paralogue, not a missed annotation; filtering those changed the refutation rate from ~75% to ~22%.
Licence and reuse
GanoDB's own derived data — gene models, annotations, predicted structures, scores and tables — is released under Creative Commons Attribution 4.0 (CC BY 4.0). You may share and adapt it, including commercially, provided you give attribution.
This licence covers GanoDB's outputs only. It does not override the terms of the upstream resources listed below, some of which restrict redistribution — where a record is derived from one of those, that resource's terms travel with it.
How to cite
Until a data-descriptor paper is published, cite the resource and the date you accessed it:
GanoDB: a genome and applied-research database for Ganoderma boninense. INBIOSIS. https://ganodb.inbiosis.org (accessed YYYY-MM-DD).
Please also cite the underlying reference assembly (GCA_002900995.3) and the specific upstream databases whose records you used — a candidate effector taken from here rests on EffectorP, and a mutant phenotype rests on PHI-base.
Sources, and their terms
GanoDB redistributes derived records from the resources below. Each is the work of another group and is credited here; follow the link for that resource's own licence before reusing its records beyond GanoDB.
| Resource | Used for |
|---|---|
| NCBI (GenBank, SRA, NR) | reference assembly, public RNA-seq and genome runs, homology |
| PHI-base | mutant phenotypes, transferred by homology |
| STRING | predicted interactions, transferred through 1:1 orthologs |
| KEGG | pathway assignment and pathway maps |
| CAZy / dbCAN | carbohydrate-active enzyme families |
| MEROPS | protease families |
| InterPro / Pfam | domains |
| eggNOG | orthology-based function, COG, GO, KO |
| MIBiG / antiSMASH | biosynthetic gene clusters |
| TCDB | transporter classification |
| PubMed | the species bibliography and gene-level references |
KEGG in particular: pathway maps are rendered by KEGG's own servers and are subject to the KEGG licence. Academic use of the KEGG website is free; other uses may require a licence from Kanehisa Laboratories.
Contact and corrections
GanoDB is developed and maintained by Nor Azlan Nor Muhammad at the Institute of Systems Biology (INBIOSIS), Universiti Kebangsaan Malaysia.
Enquiries, collaboration and corrections: norazlannm@ukm.edu.my. Corrections are actively wanted — if a gene model, an annotation or a caveat here is wrong, that is worth knowing about. Please include the identifier and the page you saw it on.
Reproducibility
Every table behind every page is downloadable from Downloads, and the provenance of each sequencing run — including which ones GanoDB actually consumed — is on the sequencing page. The pipeline that produced the database is version-controlled alongside the application.