Convolutional neural networks (CNNs) are widely used for biomedical image classification, yet it remains unclear under which conditions training on reduced subsets of available data can provide reliable guidance during model development, how much training data is required to achieve stable and comparable performance across CNN architectures, and whether increasing the training set size always leads to improved generalization or can sometimes result in degraded performance. We study this question across four biomedical datasets (breast mammography, pediatric chest X-ray, brain tumor MRI, and skin lesion dermoscopy) using the full grid of 39 CNN architectures (1–3 convolutional layers, 16/32/64 filters) from our companion architectural study, training each configuration from scratch on seven proportions of the training data (5%, 10%, 20%, 40%, 60%, 80%, and 100%) over 5 independent runs per configuration, with the test set held at a fixed size across all sample-size conditions to ensure a like-for-like comparison of generalization performance. The analysis investigates three complementary aspects of sample-size sensitivity: the stabilization of the training–test generalization gap as training-set size increases, the reliability of architecture rankings obtained from reduced training fractions as a proxy for the full-dataset ranking, and the monotonicity of test performance with respect to training-set size. Taken together, the results point to a rough four-band pattern of reliability across the sampled fractions—unstable below 20% of the training set, of uncertain overfitting status between 20% and 60% , comparatively stable between 60% and 80% , and potentially counterproductive beyond 80% —while showing that this pattern is itself dataset-dependent and offers no guarantee on architecture ranking, arguing against reduced-fraction screening as a reliable shortcut for CNN architecture selection in biomedical imaging. All code and datasets are publicly released for reproducibility.
CNN Sample-Size Effects Across Biomedical Datasets: A Reliability Pattern in Overfitting, Ranking, and Monotonicity
Sgarro, Giacinto Angelo;Grilli, Luca
2026-01-01
Abstract
Convolutional neural networks (CNNs) are widely used for biomedical image classification, yet it remains unclear under which conditions training on reduced subsets of available data can provide reliable guidance during model development, how much training data is required to achieve stable and comparable performance across CNN architectures, and whether increasing the training set size always leads to improved generalization or can sometimes result in degraded performance. We study this question across four biomedical datasets (breast mammography, pediatric chest X-ray, brain tumor MRI, and skin lesion dermoscopy) using the full grid of 39 CNN architectures (1–3 convolutional layers, 16/32/64 filters) from our companion architectural study, training each configuration from scratch on seven proportions of the training data (5%, 10%, 20%, 40%, 60%, 80%, and 100%) over 5 independent runs per configuration, with the test set held at a fixed size across all sample-size conditions to ensure a like-for-like comparison of generalization performance. The analysis investigates three complementary aspects of sample-size sensitivity: the stabilization of the training–test generalization gap as training-set size increases, the reliability of architecture rankings obtained from reduced training fractions as a proxy for the full-dataset ranking, and the monotonicity of test performance with respect to training-set size. Taken together, the results point to a rough four-band pattern of reliability across the sampled fractions—unstable below 20% of the training set, of uncertain overfitting status between 20% and 60% , comparatively stable between 60% and 80% , and potentially counterproductive beyond 80% —while showing that this pattern is itself dataset-dependent and offers no guarantee on architecture ranking, arguing against reduced-fraction screening as a reliable shortcut for CNN architecture selection in biomedical imaging. All code and datasets are publicly released for reproducibility.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


