REDIBAGG: Reducing the training set size in ensemble machine learning-based prediction models
Identificadores
URI: http://hdl.handle.net/10498/35835
DOI: 10.1016/j.engappai.2025.110382
URL: https://authors.elsevier.com/a/1klJL3OWJ9CRyl
ISSN: ISSN: 0952-1976
Ficheros
Estadísticas
Métricas y Citas
Metadatos
Mostrar el registro completo del ítemFecha
2025-06-01Departamento/s
Ingeniería InformáticaFuente
Engineering Applications of Artificial Intelligence Volume 149, 1 110382Resumen
Machine learning-based algorithms have gained wide acceptance over the years due
to their high generalization capabilities for a wide range of classification applications.
Although these algorithms demonstrate potential and promising performance, they are
often limited by speed, particularly when training on large databases. Big data sets with many instances have storage requirements and execution times that can be excessive. This study proposes reducing the sample size generated by the bootstrap resampling method in Ensemble Machine Learning-based models and evaluates their generalization capability on unknown data. Reduced bootstrap samples are employed in the training phase of the Bagging ensemble model. This approach reduces execution times and, consequently, storage requirements. The proposed method was tested on classification tasks, effectively reducing training subset size without compromising performance. Experimental results demonstrate that this approach achieves execution times reductions of up to 70% for some data sets.
This reduction has no impact on accuracy, whereas maintaining levels comparable to classical Bagging and its variants. On average, the training subset size was reduced by 25% compared to the original size.
Materias
Artificial intelligence; Machine Learning; Big data; Bagging ensemble; Bootstrap; Reduced sampleColecciones
- Artículos Científicos [11777]






