Trustworthy AI for Neurodegenerative Disease Screening: Explainability, Clinical Accountability, and Human–AI Collaboration
DOI:
https://doi.org/10.63646/BCLE3872Keywords:
Trustworthy AI; neurodegenerative disease; explainable AI; clinical decision support; accountability; human–AI collaboration; medical imagingAbstract
Neurodegenerative diseases such as Parkinson's disease and Alzheimer's disease are placing a growing burden on ageing societies, while traditional clinical screening still depends heavily on subjective expert evaluation and long diagnostic pathways. Recent deep-learning models on resting-state functional MRI and other neuroimaging data have improved discrimination accuracy, but their adoption in clinical practice remains limited because they offer little insight into how a decision is reached, who is accountable when it is wrong, and how it should be combined with the judgement of an experienced clinician. This article proposes a sociotechnical framework for trustworthy AI in neurodegenerative disease screening that integrates three pillars — explainability, clinical accountability, and human–AI collaboration — and connects them into a single deployable workflow. We characterise faithful explanation methods for graph-based brain-network models, including saliency, concept-based attribution, subgraph rationale, and uncertainty quantification, and discuss the role of regulatory frameworks (FDA SaMD, EU AI Act, NMPA) in defining audit trails, post-market surveillance, and liability allocation. We then formalise a collaborative diagnostic workflow in which clinicians retain final authority while AI provides risk scores, calibrated confidence, and structured rationale, and we describe a closed feedback loop that links clinical override events to bias monitoring and model retraining. An empirical evaluation across four hospital cohorts (640 participants) shows that the framework preserves screening performance (pooled AUROC 0.881) while improving clinician trust calibration and reducing decision time. The findings suggest that trustworthy AI for neurodegenerative screening is achievable only when explainability, accountability, and collaboration are designed jointly rather than treated as independent technical add-ons.
How to Cite
References
Adadi, A., & Berrada, M. (2018). Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access, 6, 52138–52160. https://doi.org/10.1109/ACCESS.2018.2870052
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., & Kim, B. (2018). Sanity checks for saliency maps. Advances in Neural Information Processing Systems, 31, 9525–9536. https://doi.org/10.48550/arXiv.1810.03292
Asan, O., Bayrak, A. E., & Choudhury, A. (2020). Artificial intelligence and human trust in healthcare: Focus on clinicians. Journal of Medical Internet Research, 22(6), e15154. https://doi.org/10.2196/15154
Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, Article 81. https://doi.org/10.1145/3411764.3445717
Benjamens, S., Dhunnoo, P., & Meskó, B. (2020). The state of artificial intelligence-based FDA-approved medical devices and algorithms: An online database. npj Digital Medicine, 3(1), 118. https://doi.org/10.1038/s41746-020-00324-0
Bi, X.-A., Hu, X., Wu, H., & Wang, Y. (2020). Multimodal data analysis of Alzheimer's disease based on clustering evolutionary random forest. IEEE Journal of Biomedical and Health Informatics, 24(10), 2973–2983. https://doi.org/10.1109/JBHI.2020.2973324
Cai, C. J., Reif, E., Hegde, N., Hipp, J., Kim, B., Smilkov, D., Wattenberg, M., Viegas, F., Corrado, G. S., Stumpe, M. C., & Terry, M. (2019). Human-centered tools for coping with imperfect algorithms during medical decision-making. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, Paper 4. https://doi.org/10.1145/3290605.3300234
Char, D. S., Shah, N. H., & Magnus, D. (2018). Implementing machine learning in health care — Addressing ethical challenges. New England Journal of Medicine, 378(11), 981–983. https://doi.org/10.1056/NEJMp1714229
Chen, I. Y., Pierson, E., Rose, S., Joshi, S., Ferryman, K., & Ghassemi, M. (2021). Ethical machine learning in healthcare. Annual Review of Biomedical Data Science, 4, 123–144. https://doi.org/10.1146/annurev-biodatasci-092820-114757
Cohen, I. G., Amarasingham, R., Shah, A., Xie, B., & Lo, B. (2014). The legal and ethical concerns that arise from using complex predictive analytics in health care. Health Affairs, 33(7), 1139–1147. https://doi.org/10.1377/hlthaff.2014.0048
Doshi-Velez, F., & Kim, B. (2017). Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608. https://doi.org/10.48550/arXiv.1702.08608
European Commission. (2021). Proposal for a regulation laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). COM(2021) 206 final. EUR-Lex: CELEX:52021PC0206.
Filippi, M., Sarasso, E., & Agosta, F. (2019). Resting-state functional MRI in Parkinsonian syndromes. Movement Disorders Clinical Practice, 6(2), 104–117. https://doi.org/10.1002/mdc3.12730
Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V., Luetge, C., Madelin, R., Pagallo, U., Rossi, F., Schafer, B., Valcke, P., & Vayena, E. (2018). AI4People — An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Minds and Machines, 28(4), 689–707. https://doi.org/10.1007/s11023-018-9482-5
Gal, Y., & Ghahramani, Z. (2016). Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. Proceedings of the 33rd International Conference on Machine Learning, 48, 1050–1059. https://doi.org/10.48550/arXiv.1506.02142
Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127. https://doi.org/10.1136/amiajnl-2011-000089
Hooker, S., Erhan, D., Kindermans, P.-J., & Kim, B. (2019). A benchmark for interpretability methods in deep neural networks. Advances in Neural Information Processing Systems, 32, 9737–9748. https://doi.org/10.48550/arXiv.1806.10758
Jack, C. R., Bennett, D. A., Blennow, K., Carrillo, M. C., Dunn, B., Haeberlein, S. B., Holtzman, D. M., Jagust, W., Jessen, F., Karlawish, J., Liu, E., Molinuevo, J. L., Montine, T., Phelps, C., Rankin, K. P., Rowe, C. C., Scheltens, P., Siemers, E., Snyder, H. M., & Sperling, R. (2018). NIA-AA Research Framework: Toward a biological definition of Alzheimer's disease. Alzheimer's & Dementia, 14(4), 535–562. https://doi.org/10.1016/j.jalz.2018.02.018
Jacovi, A., & Goldberg, Y. (2020). Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4198–4205. https://doi.org/10.18653/v1/2020.acl-main.386
Kendall, A., & Gal, Y. (2017). What uncertainties do we need in Bayesian deep learning for computer vision? Advances in Neural Information Processing Systems, 30, 5574–5584. https://doi.org/10.48550/arXiv.1703.04977
Khosla, M., Jamison, K., Ngo, G. H., Kuceyeski, A., & Sabuncu, M. R. (2019). Machine learning in resting-state fMRI analysis. Magnetic Resonance Imaging, 64, 101–121. https://doi.org/10.1016/j.mri.2019.05.031
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., & Sayres, R. (2018). Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). Proceedings of the 35th International Conference on Machine Learning, 80, 2668–2677. https://doi.org/10.48550/arXiv.1711.11279
Kim, B.-H., Ye, J. C., & Kim, J.-J. (2021). Learning dynamic graph representation of brain connectome with spatio-temporal attention. Advances in Neural Information Processing Systems, 34, 4314–4327. https://doi.org/10.48550/arXiv.2105.13495
Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., & Liang, P. (2020). Concept bottleneck models. Proceedings of the 37th International Conference on Machine Learning, 119, 5338–5348. https://doi.org/10.48550/arXiv.2007.04612
Lakshminarayanan, B., Pritzel, A., & Blundell, C. (2017). Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in Neural Information Processing Systems, 30, 6402–6413. https://doi.org/10.48550/arXiv.1612.01474
Larson, D. B., Magnus, D. C., Lungren, M. P., Shah, N. H., & Langlotz, C. P. (2020). Ethics of using and sharing clinical imaging data for artificial intelligence: A proposed framework. Radiology, 295(3), 675–682. https://doi.org/10.1148/radiol.2020192536
Li, X., Zhou, Y., Dvornek, N., Zhang, M., Gao, S., Zhuang, J., Scheinost, D., Staib, L. H., Ventola, P., & Duncan, J. S. (2021). BrainGNN: Interpretable brain graph neural network for fMRI analysis. Medical Image Analysis, 74, 102233. https://doi.org/10.1016/j.media.2021.102233
Lu, Y. (2017). Industry 4.0: A survey on technologies, applications and open research issues. Journal of Industrial Information Integration, 6, 1–10. https://doi.org/10.1016/j.jii.2017.04.005
Lu, Y. (2019). Artificial intelligence: A survey on evolution, models, applications and future trends. Journal of Management Analytics, 6(1), 1–29. https://doi.org/10.1080/23270012.2019.1570365
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774. https://doi.org/10.48550/arXiv.1705.07874
Luo, D., Cheng, W., Xu, D., Yu, W., Zong, B., Chen, H., & Zhang, X. (2020). Parameterized explainer for graph neural network. Advances in Neural Information Processing Systems, 33, 19620–19631. https://doi.org/10.48550/arXiv.2011.04573
Maassen, O., Fritsch, S., Palm, J., Deffge, S., Kunze, J., Marx, G., Riedel, M., Schuppert, A., & Bickenbach, J. (2021). Future medical artificial intelligence application requirements and expectations of physicians in German university hospitals: Web-based survey. Journal of Medical Internet Research, 23(3), e26646. https://doi.org/10.2196/26646
Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S., & Floridi, L. (2016). The ethics of algorithms: Mapping the debate. Big Data & Society, 3(2), 205395171667967. https://doi.org/10.1177/2053951716679679
Muehlematter, U. J., Daniore, P., & Vokinger, K. N. (2021). Approval of artificial intelligence and machine learning-based medical devices in the USA and Europe (2015–20): A comparative analysis. The Lancet Digital Health, 3(3), e195–e203. https://doi.org/10.1016/S2589-7500(20)30292-2
NMPA. (2022). Guideline on registration review of artificial intelligence medical devices. Beijing: Center for Medical Device Evaluation, National Medical Products Administration.
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342
Postuma, R. B., & Berg, D. (2019). Prodromal Parkinson's disease: The decade past, the decade to come. Movement Disorders, 34(5), 665–675. https://doi.org/10.1002/mds.27670
Postuma, R. B., Berg, D., Stern, M., Poewe, W., Olanow, C. W., Oertel, W., Obeso, J., Marek, K., Litvan, I., Lang, A. E., Halliday, G., Goetz, C. G., Gasser, T., Dubois, B., Chan, P., Bloem, B. R., Adler, C. H., & Deuschl, G. (2015). MDS clinical diagnostic criteria for Parkinson's disease. Movement Disorders, 30(12), 1591–1601. https://doi.org/10.1002/mds.26424
Price, W. N., Gerke, S., & Cohen, I. G. (2019). Potential liability for physicians using artificial intelligence. JAMA, 322(18), 1765–1766. https://doi.org/10.1001/jama.2019.15064
Reddy, S., Allan, S., Coghlan, S., & Cooper, P. (2020). A governance model for the application of AI in health care. Journal of the American Medical Informatics Association, 27(3), 491–497. https://doi.org/10.1093/jamia/ocz192
Rizzo, G., Copetti, M., Arcuti, S., Martino, D., Fontana, A., & Logroscino, G. (2016). Accuracy of clinical diagnosis of Parkinson disease: A systematic review and meta-analysis. Neurology, 86(6), 566–576. https://doi.org/10.1212/WNL.0000000000002350
Strickland, E. (2019). IBM Watson, heal thyself: How IBM overpromised and underdelivered on AI health care. IEEE Spectrum, 56(4), 24–31. https://doi.org/10.1109/MSPEC.2019.8678513
Sullivan, H. R., & Schweikart, S. J. (2019). Are current tort liability doctrines adequate for addressing injury caused by AI? AMA Journal of Ethics, 21(2), E160–166. https://doi.org/10.1001/amajethics.2019.160
Sundararajan, M., Taly, A., & Yan, Q. (2017). Axiomatic attribution for deep networks. Proceedings of the 34th International Conference on Machine Learning, 70, 3319–3328. https://doi.org/10.48550/arXiv.1703.01365
Tonekaboni, S., Joshi, S., McCradden, M. D., & Goldenberg, A. (2019). What clinicians want: Contextualizing explainable machine learning for clinical end use. Proceedings of the 4th Machine Learning for Healthcare Conference, 106, 359–380. https://doi.org/10.48550/arXiv.1905.05134
U.S. FDA. (2021). Artificial intelligence/machine learning (AI/ML)-based software as a medical device (SaMD) action plan. Silver Spring, MD: U.S. Food and Drug Administration. https://www.fda.gov/media/145022/download
WHO. (2021). Ethics and governance of artificial intelligence for health: WHO guidance. Geneva: World Health Organization. ISBN 978-92-4-002920-0.
Ying, R., Bourgeois, D., You, J., Zitnik, M., & Leskovec, J. (2019). GNNExplainer: Generating explanations for graph neural networks. Advances in Neural Information Processing Systems, 32, 9240–9251. https://doi.org/10.48550/arXiv.1903.03894
Zhang, C., & Lu, Y. (2021). Study on artificial intelligence: The state of the art and future prospects. Journal of Industrial Information Integration, 23, 100224. https://doi.org/10.1016/j.jii.2021.100224
van Leeuwen, K. G., Schalekamp, S., Rutten, M. J. C. M., van Ginneken, B., & de Rooij, M. (2021). Artificial intelligence in radiology: 100 commercially available products and their scientific evidence. European Radiology, 31(6), 3797–3804. https://doi.org/10.1007/s00330-021-07892-z