Method & references
The identification matrix
The knowledge base is a matrix of n taxa × m characters (species). Each cell is the percentage of samples of that taxon in which the character is present (the “percent positive” value). Taxa may be depositional environments, foram bands or pollen zones, depending on the matrix used.
Willcox probability
For a sample, every character is scored present (1) or absent (0). The likelihood of the sample for a given taxon is the product over all characters of the taxon’s percent-positive value for present characters and its complement (100 − value) for absent characters, after division by 100. The likelihoods are normalised over all taxa so that they sum to one; these normalised values are the Willcox probabilities, and the three highest are reported for each sample.
Calculations are performed in logarithmic space, which removes the risk of numerical underflow on long species lists; the results are identical to the original algorithm to machine precision. An optional “positive entries only” mode computes the likelihood from present characters alone.
Diagnostics and diversity
- Species against: characters whose expected frequency in a taxon differs from the sample by more than 90% — the species that argue against (or conspicuously for) each of the three best identifications.
- Diversity indices (quantitative samples): total specimens, planktonic/benthonic ratio, Yule–Simpson index and Fisher’s alpha.
History
The method derives from Sneath (1979). The original program, BULKMAT, was written by P. Lesslar (XGS/1, October 1984) for the identification of well samples using presence–absence data. This web edition reproduces its calculations and report format exactly, and adds a searchable species list, on-page options and structured logging of submissions for the further development of the identification knowledge base.
References
- Sneath, P.H.A. (1979). BASIC program for identification of an unknown with presence–absence data against an identification matrix of percent positive characters. Computers & Geosciences 5, 195–213.
- Willcox, W.B., Lapage, S.P., Bascomb, S. & Curtis, M.A. (1973). Identification of bacteria by computer: theory and programming. Journal of General Microbiology 77, 317–330.
