YaBeSH Engineering and Technology Library

    • Journals
    • PaperQuest
    • YSE Standards
    • YaBeSH
    • Login
    View Item 
    •   YE&T Library
    • ASME
    • Journal of Engineering and Science in Medical Diagnostics and Therapy
    • View Item
    •   YE&T Library
    • ASME
    • Journal of Engineering and Science in Medical Diagnostics and Therapy
    • View Item
    • All Fields
    • Source Title
    • Year
    • Publisher
    • Title
    • Subject
    • Author
    • DOI
    • ISBN
    Advanced Search
    JavaScript is disabled for your browser. Some features of this site may not work without it.

    Archive

    Reliable Benchmarking of Breast Ultrasound Lesion Classification Requires Patient-Level Validation

    Source: Journal of Engineering and Science in Medical Diagnostics and Therapy:;2026:;volume( 009 ):;issue:004::page 209
    Author:
    Wang, Lulu
    DOI: 10.1115/1.4071712
    Publisher: The American Society of Mechanical Engineers (ASME)
    Abstract: Abstract. Breast ultrasound is widely used for lesion characterization, yet reported deep-learning performance varies substantially with dataset composition, preprocessing, and validation design. A major source of bias arises when patient-level separation is not enforced, allowing correlated images from the same subject to inflate performance estimates. This study presents a leakage-aware benchmark of convolutional neural networks (CNNs), a Vision Transformer (ViT), and a CNN–transformer late-fusion configuration for benign-versus-malignant breast ultrasound classification on the BUS-UCLM dataset. After exclusion of normal-category images, the final cohort comprised 264 images from 36 patients, including 174 benign and 90 malignant images. Seven CNN-family models, one ViT baseline, and one ResNet18–ViT probability-level late-fusion configuration were evaluated using strict patient-level fivefold cross-validation. Additional analyses included fusion ablation, paired Wilcoxon signed-rank testing, gradient-weighted class activation mapping (Grad-CAM) visualization, and an image-level leakage demonstration. Under strict patient-level evaluation, performance was moderate across all models. GoogLeNet achieved the highest mean accuracy (61.48%), InceptionV3 achieved the highest mean macro-F1 (59.24%), and ResNet50 achieved the highest mean area under the receiver operating characteristic curve (AUC) (0.6699), whereas the standalone ViT showed weaker overall discrimination. The late-fusion configuration remained competitive in threshold-dependent metrics but did not surpass the strongest CNN baselines in AUC. Overall, no architecture demonstrated a clear advantage across both threshold-dependent and threshold-independent metrics. By contrast, image-level splitting substantially inflated apparent performance, underscoring the importance of rigorous patient-level separation for credible benchmarking in breast ultrasound classification.
    • Download: (1.418Mb)
    • Show Full MetaData Hide Full MetaData
    • Get RIS
    • Item Order
    • Go To Publisher
    • Statistics

      Reliable Benchmarking of Breast Ultrasound Lesion Classification Requires Patient-Level Validation

    URI
    https://yetl.yabesh.ir/yetl1/handle/yetl/4316011
    Collections
    • Journal of Engineering and Science in Medical Diagnostics and Therapy

    Show full item record

    contributor authorWang, Lulu
    date accessioned2026-08-23T08:03:17Z
    date available2026-08-23T08:03:17Z
    date copyright2026/11/01
    date issued2026
    identifier issn2572-7958
    identifier otherjesmdt-26-1007.pdf
    identifier urihttp://yetl.yabesh.ir/yetl1/handle/yetl/4316011
    description abstractAbstract. Breast ultrasound is widely used for lesion characterization, yet reported deep-learning performance varies substantially with dataset composition, preprocessing, and validation design. A major source of bias arises when patient-level separation is not enforced, allowing correlated images from the same subject to inflate performance estimates. This study presents a leakage-aware benchmark of convolutional neural networks (CNNs), a Vision Transformer (ViT), and a CNN–transformer late-fusion configuration for benign-versus-malignant breast ultrasound classification on the BUS-UCLM dataset. After exclusion of normal-category images, the final cohort comprised 264 images from 36 patients, including 174 benign and 90 malignant images. Seven CNN-family models, one ViT baseline, and one ResNet18–ViT probability-level late-fusion configuration were evaluated using strict patient-level fivefold cross-validation. Additional analyses included fusion ablation, paired Wilcoxon signed-rank testing, gradient-weighted class activation mapping (Grad-CAM) visualization, and an image-level leakage demonstration. Under strict patient-level evaluation, performance was moderate across all models. GoogLeNet achieved the highest mean accuracy (61.48%), InceptionV3 achieved the highest mean macro-F1 (59.24%), and ResNet50 achieved the highest mean area under the receiver operating characteristic curve (AUC) (0.6699), whereas the standalone ViT showed weaker overall discrimination. The late-fusion configuration remained competitive in threshold-dependent metrics but did not surpass the strongest CNN baselines in AUC. Overall, no architecture demonstrated a clear advantage across both threshold-dependent and threshold-independent metrics. By contrast, image-level splitting substantially inflated apparent performance, underscoring the importance of rigorous patient-level separation for credible benchmarking in breast ultrasound classification.
    publisherThe American Society of Mechanical Engineers (ASME)
    titleReliable Benchmarking of Breast Ultrasound Lesion Classification Requires Patient-Level Validation
    typeJournal Paper
    journal volume9
    journal issue4
    journal titleJournal of Engineering and Science in Medical Diagnostics and Therapy
    identifier doi10.1115/1.4071712
    journal fristpage209
    journal lastpage249
    page41
    treeJournal of Engineering and Science in Medical Diagnostics and Therapy:;2026:;volume( 009 ):;issue:004
    contenttypeFulltext
    DSpace software copyright © 2002-2015  DuraSpace
    نرم افزار کتابخانه دیجیتال "دی اسپیس" فارسی شده توسط یابش برای کتابخانه های ایرانی | تماس با یابش
    yabeshDSpacePersian
     
    DSpace software copyright © 2002-2015  DuraSpace
    نرم افزار کتابخانه دیجیتال "دی اسپیس" فارسی شده توسط یابش برای کتابخانه های ایرانی | تماس با یابش
    yabeshDSpacePersian