Aprendizagem de Classificadores para Identificação de Bactérias: relação entre as medidas de complexidade de dados e o desempenho dos classificadores
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Universidade Estadual de Ponta Grossa
Abstract
In the agricultural environment, some bacteria have been used as active in biocontrol and plant
growth. This has motivated the development of software tools to automatically detect their presence
in soil samples. One way to proceed with this identification is the development of classifiers
that use MALDI / TOF mass spectra patterns to check the frequency of certain ribosomal
proteins in the sample. The selection of a classification function that fits the target problem has
a great influence on the classifier’s performance, this has encouraged the use of scores, called
data complexity measures. Such scores describe certain characteristics of the database and may
provide support for choosing the classification function. During the process of generating data
from mass spectrometry, it is common for data to be unbalanced, which adversely affects the
data complexity measures. Considering the above, this work applies an experimental protocol
to verify the influence of unbalanced data on the performance of classifiers and on complexity
measures. The classifying models used in the experiments were logistic regression and QDA,
which were trained to identify bacteria of the genera Bacillus and Rhizobium. The performance
of the classifiers showed a strong to moderate relationship with the unbalanced data problem.
Two data complexity indexes, L2B and N3B, have been proposed and submitted to tests along
with the indexes found in the literature. The results show that the measures F3, Density, N3B
and L2B are related to the performance of the classifiers trained with unbalanced data. Such
measures were evaluated for their ability to predict the balanced accuracy of the models. When
identifying bacteria of the genera Bacillus, the measure of best relation to the performance of
the models was the N3B measure. In the case of the identification of the genera Rhizobium, the
measure of best association with the logistic model was L2B and N3B for the quadratic model.
Description
Citation
FEDACZ, Gabriel Lucas. Aprendizagem de Classificadores para Identificação de Bactérias: relação entre as medidas de complexidade de dados e o desempenho dos classificadores. 2020. Dissertação (Mestrado em Computação Aplicada) - Universidade Estadual de Ponta Grossa, Ponta Grossa, 2020.
Endorsement
Review
Supplemented By
Referenced By
Creative Commons license
Except where otherwised noted, this item's license is described as Acesso Aberto
