lnu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Improving Classification in Imbalanced Educational Datasets using Over-sampling
Linnaeus University, Faculty of Technology, Department of computer science and media technology (CM). Linnaeus University, Linnaeus Knowledge Environments, Digital Transformations. (EdTechLnu)ORCID iD: 0000-0002-3297-0189
Linnaeus University, Faculty of Technology, Department of computer science and media technology (CM). Linnaeus University, Linnaeus Knowledge Environments, Digital Transformations.ORCID iD: 0000-0002-2901-935X
Linnaeus University, Faculty of Technology, Department of computer science and media technology (CM). Linnaeus University, Linnaeus Knowledge Environments, Digital Transformations.ORCID iD: 0000-0002-6937-345X
Linnaeus University, Faculty of Social Sciences, Department of Pedagogy and Learning. Linnaeus University, Linnaeus Knowledge Environments, Digital Transformations.ORCID iD: 0000-0002-3738-7945
2020 (English)In: Proceedings of the 28th international conference on computer in education, Asia-Pacific Society for Computers in Education, 2020, Vol. 1, p. 278-283Conference paper, Published paper (Refereed)
Abstract [en]

Learning Analytics (LA) involves a growing range of methods for understanding and optimizing learning and the environments in which it occurs. Different Machine Learning (ML) algorithms or learning classifiers can be used to implement LA, with the goal of predicting learning outcomes and classifying the data into predetermined categories. Many educational datasets are imbalanced, where the number of samples in one category is significantly larger than in other categories. Ordinarily, it is ML’s performance on the minority categories that is the most important. Since most ML classification algorithms ignore the minority categories, and in turn have poor performance, so learning from imbalanced datasets is really challenging. In order to address this challenge and also to improve the performance of different classifiers, Synthetic Minority Over-sampling Technique (SMOTE) is used to oversample the minority categories. In this paper, the accuracy of seven well-known classifiers considering 5 and 10-fold cross-validation and the F1-score are compared. The imbalanced dataset collected based on self-regulated learning activities contains the learning behaviour of 6,423 medical students who used a web-based study platform—Hypocampus—with different educational topics for one year. Also, two diagnostic tools including Area Under the Receiver Operating Characteristics (AUC-ROC) curves and Precision-Recall (PR) curves are applied to predict probabilities of an observation belonging to each category in a classification problem. Using these diagnostic tools may help LA researchers on how to deal with imbalanced educational datasets. The outcomes of our experimental results show that Neural Network with 92.77% in 5-fold cross-validation, 93.20% in 10-fold cross-validation and 0.95 in F1-score has the highest accuracy and performance compared to other classifiers when we applied the SMOTE technique. Also, the probability of detection in different classifiers using SMOTE has shown a significant improvement. 

Place, publisher, year, edition, pages
Asia-Pacific Society for Computers in Education, 2020. Vol. 1, p. 278-283
Keywords [en]
Learning Analytics, Imbalanced Dataset, Machine Learning, SMOTE, ROC, PR
National Category
Computer Sciences
Research subject
Computer and Information Sciences Computer Science, Computer Science; Computer and Information Sciences Computer Science
Identifiers
URN: urn:nbn:se:lnu:diva-101549Scopus ID: 2-s2.0-85099475207ISBN: 9789869721455 (electronic)OAI: oai:DiVA.org:lnu-101549DiVA, id: diva2:1535608
Conference
28th International Conference on Computers in Education
Available from: 2021-03-09 Created: 2021-03-09 Last updated: 2026-04-16Bibliographically approved

Open Access in DiVA

fulltext(1344 kB)475 downloads
File information
File name FULLTEXT02.pdfFile size 1344 kBChecksum SHA-512
62097e2ad73684a402e6cc41ca7b132eb92fc245c01b60f6b279548f19d383bdaa5fde0e520dcc9b30af446f4dafe95c19a166b2a7f6e90c0f4bf0d1706fe279
Type fulltextMimetype application/pdf

Scopus

Authority records

Mohseni, ZeynabMartins, Rafael MessiasMilrad, MarceloMasiello, Italo

Search in DiVA

By author/editor
Mohseni, ZeynabMartins, Rafael MessiasMilrad, MarceloMasiello, Italo
By organisation
Department of computer science and media technology (CM)Digital TransformationsDepartment of Pedagogy and Learning
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 475 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

isbn
urn-nbn

Altmetric score

isbn
urn-nbn
Total: 857 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf