Multi-architecture deep learning framework for glaucoma detection using CNN and vision transformer models

Citation

Abstract

Glaucoma as a primary cause of permanent blindness is still unnoticeable in the initial stages due to the absence of symptoms. The purpose of this study is to design an automated system of glaucoma detection, which will be based on the use of deep learning algorithms to retinal fundus images. The proposed work has introduced a new image preprocessing methodology that will involve the combination of Contrast Limited Adaptive Histogram Equalization (CLAHE), gamma correction, and green channel extraction in order to enhance the quality of the image and make the bright points including optic nerve head and retinal vasculature more discernible and visible that is vital in the detection of glaucoma correctly. The efficacy of four various deep learning structures, such as ResNet50, EfficientNetB0, SwinTransformer, and Graph Convolutional Networks (GCNs) in glaucoma detection, have been tested. In the present study, the performance of four deep learning models, namely ResNet50, EfficientNetB0, SwinTransformer, and Graph Convolutional Networks (GCNs), is evaluated to determine the effectiveness of the models in the detection of glaucoma. Among the four models, the ResNet50 model recorded the highest performance, with an accuracy of 99.47%, sensitivity of 100%, and specificity of 99.33%. In the present study, the application of the ResNet50 model in the clinical domain is also explored. For the application, the model is converted into the pytorch format, which is compatible with all platforms. It is also optimized to run the model efficiently. A streamlit interface is developed to deploy the ResNet50 model in the clinical domain for the detection of glaucoma. In the Gradio interface, the clinicians can upload the images of the retinal fundus, and the model will give instant results. In the proposed system, the cloud, edge, and hybrid architectures are used. Nevertheless, the nextgeneration study must focus on the multi-center validation, which will assess the generalizability of the offered model. The applicability and credibility of the proposed model will be enhanced by the addition of other diagnostic tools, including Optical Coherence Tomography (OCT), and the development of eXplainable Artificial Intelligence (XAI) models. The suggested study demonstrates the feasibility of deep learning application in the glaucoma detection, and it offers an effective solution to the early diagnosis and treatment of the disease.

Description

This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2026.
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 72-75).

Publisher Link

Type

Thesis

Creative Commons license

Attribution-NonCommercial-NoDerivatives 4.0 International

Except where otherwise noted, this item's license is described as

Attribution-NonCommercial-NoDerivatives 4.0 International