RESEARCH / DISCOVERY
← Back to the library

Πρόβλεψη της νόσου Parkinson με τη χρήση μεθόδων μηχανικής μάθησης σε φωνητικά δεδομένα

Πρόβλεψη της νόσου Parkinson με τη χρήση μεθόδων μηχανικής μάθησης σε φωνητικά δεδομένα

Read the original publication

Where did the research take place?

The study site has not been established. Author addresses may differ from where the research occurred.

Explore research worldwide

A plain-language reading has not been prepared for this paper yet.

Original abstract

Η παρούσα διπλωματική εργασία ασχολείται με την πρόβλεψη της εμφάνισης της νόσου του Parkinson σε διάφορα άτομα μέσω της εφαρμογής μεθόδων μηχανικής μάθησης σε ένα σύνολο φωνητικών δεδομένων με 24 μεταβλητές και 195 εγγραφές. Στο πρώτο κεφάλαιο, αναφερόμαστε στο σκοπό και τους στόχους της εργασίας και μιλούμε για τη νόσο του Parkinson. Επιπροσθέτως, παραθέτουμε μία βιβλιογραφική επισκόπηση σε σχετικές έρευνες που παρακίνησαν τη μελέτη μας και περιγράφουμε τη μεθοδολογία που εφαρμόζουμε, επισημαίνοντας και διαφορές σε αυτή συγκριτικά με τη βιβλιογραφία. Στο κλείσιμο του πρώτου κεφαλαίου, συζητούμε σχετικά με τη μηχανική μάθηση. Στο δεύτερο κεφάλαιο, αρχικά εστιάζουμε στην προέλευση, τις μεταβλητές και ορισμένα δομικά χαρακτηριστικά του συνόλου δεδομένων. Έπειτα, διερευνούμε τόσο τη συνολική κατανομή των μεταβλητών του όσο και τη σχετική κατανομή των ανεξάρτητων μεταβλητών του (φωνητικά χαρακτηριστικά) σε σχέση με την εξαρτημένη (υποδεικνύει αν κάποιος είναι υγιής ή άρρωστος). Στο τέλος αυτού του κεφαλαίου, εξετάζουμε τις συσχετίσεις μεταξύ όλων των μεταβλητών. Στο τρίτο κεφάλαιο, αναλύουμε την προτεινόμενη μεθοδολογία και παραθέτουμε σχετικά τμήματα κώδικα υλοποίησης στη γλώσσα προγραμματισμού Python. Η μεθοδολογία ξεκινά με την αφαίρεση ασήμαντων μεταβλητών και τη διάκριση των υπολοίπων σε ανεξάρτητες (φωνητικές) και εξαρτημένες. Έπειτα διαχωρίζουμε τα δεδομένα σε σύνολα εκπαίδευσης και δοκιμής και επιλέγουμε τα πιο σημαντικά ανεξάρτητα γνωρίσματα με χρήση δέντρου απόφασης. Στη συνέχεια, φέρνουμε όλα τα φωνητικά δεδομένα σε παρόμοια κλίμακα και εξισορροπούμε τις δύο κλάσεις της μεταβλητής απόφασης στο σύνολο εκπαίδευσης. Οι αλγόριθμοι ταξινόμησης, οι οποίοι εφαρμόζονται στο τελευταίο στάδιο της μεθοδολογίας, είναι αυτοί των Κ-κοντινότερων γειτόνων, των μηχανών διανυσμάτων υποστήριξης και του τυχαίου δάσους. Στους αλγορίθμους αυτούς, όπως και στο δέντρο απόφασης, βελτιστοποιούμε ορισμένες υπερπαραμέτρους με σκοπό τη μεγιστοποίηση της μετρικής ανάκλησης. Τέλος, στο τέταρτο κεφάλαιο, αξιολογούμε τους αλγορίθμους ταξινόμησης που εφαρμόσαμε και παραθέτουμε τα συμπεράσματά μας. Όσον αφορά την απόδοση στο σύνολο δοκιμής, οι μηχανές διανυσμάτων υποστήριξης και το τυχαίο δάσος πέτυχαν την καλύτερη βαθμολογία ανάκλησης με 0.977. Το τυχαίο δάσος διέπρεψε επίσης σε ορθότητα και ακρίβεια με 0.949 και 0.956 αντίστοιχα. This thesis deals with the prediction of the occurrence of Parkinson's disease in various individuals through the application of machine learning methods to a voice data set with 24 variables and 195 recordings. In the first chapter, we refer to the purpose and the objectives of the work and we are talking about Parkinson's disease. In addition, we provide a literature review on related research that motivated our study and we describe the methodology that we apply, also highlighting differences in it compared to the literature. At the close of the first chapter, we discuss about machine learning. In the second chapter, we initially focus on the origin, the variables and certain structural characteristics of the dataset. Then, we investigate both the overall distribution of its variables and the relative distribution of its independent variables (vocal features) with respect to the dependent one (indicates if someone is healthy or diseased). In the end of this chapter, we examine the correlations between all variables. In the third chapter, we analyze the proposed methodology and list relevant code sections of the implementation in the Python programming language. The methodology starts by removing insignificant variables and distinguishing the remaining ones into independent (vocal) and dependent ones. We then separate the data into training and testing sets and select the most important independent features using a decision tree.We then bring all voice data to a similar scale and balance the two classes of the decision variable in the train set. The classification algorithms, which are applied in the last stage of the methodology, are those of K-nearest neighbors, support vector machines, and random forest. In these algorithms, as in the decision tree, we optimize some hyperparameters in order to maximize the recall metric. Finally, in the fourth chapter, we evaluate the classification algorithms that we implemented and we present our conclusions. Regarding performance on the test set, support vector machines and random forest achieved the best recall score with 0.977. Random forest also excelled in accuracy and precision with 0.949 and 0.956 respectively.

Explore another example or bring your own paper

Pasted text and PDF extraction stay on this computer. The local guide explains terms and surfaces passages; rewriting requires a configured local model. Scanned PDFs need OCR first.

RECORD & PROVENANCE