lnu.sePublications
Change search
Link to record
Permanent link

Direct link
Publications (10 of 76) Show all publications
Sylejmani, K., Bylygbashi, S., Lajçi, U. & Kastrati, Z. (2026). A curated dataset of multi-channel TV program schedules for optimization and benchmarking. Data in Brief, 65, Article ID 112568.
Open this publication in new window or tab >>A curated dataset of multi-channel TV program schedules for optimization and benchmarking
2026 (English)In: Data in Brief, E-ISSN 2352-3409, Vol. 65, article id 112568Article in journal (Refereed) Published
Abstract [en]

This article presents a collection of 515 Smart TV scheduling instances from three public sources: IPTV EPG broadcast listings, EPG.PW program guides, and upcoming YouTube livestream schedules obtained through the YouTube Data API. The data were collected and processed using automated Python scripts that download program information, convert it into a standard format, and remove entries with missing or invalid time information. All times are recorded as minutes from the start of each scheduling period to ensure consistency across different sources.

Each instance is stored as a JSON file with a consistent structure. Each file contains the scheduling period (start and end times), channel information, program time slots that don’t overlap, program categories, and quality scores from 0 to 100 for each program. Instances also include two types of constraints: strict rules that specify which channels must be available during certain time windows, and flexible preferences that assign bonus points to programs airing at preferred times. All files are organized by data source and automatically checked for correctness.

The dataset can be used for benchmarking optimization and constraint-based scheduling methods, comparing objective formulations with switching and termination penalties, testing algorithms under time-dependent feasibility constraints, and extracting instance features for dataset characterization and dimensionality-reduction workflows.

Place, publisher, year, edition, pages
Elsevier, 2026
Keywords
Real-world data, Broadcast planning, Constraint satisfaction, Time windows, Penalty functions, Genre composition, Metaheuristic search, Dimensionality reduction
National Category
Computer Sciences
Research subject
Computer and Information Sciences Computer Science
Identifiers
urn:nbn:se:lnu:diva-145139 (URN)10.1016/j.dib.2026.112568 (DOI)001694964900001 ()41732355 (PubMedID)2-s2.0-105030248890 (Scopus ID)
Available from: 2026-02-16 Created: 2026-02-16 Last updated: 2026-04-16Bibliographically approved
Mughal, N., Imran, A. S., Daudpota, S. M., Kastrati, Z. & Noor, W. (2026). Exploring potential of large language models for automated essay scoring in education. Discover Artificial Intelligence, 6, Article ID 166.
Open this publication in new window or tab >>Exploring potential of large language models for automated essay scoring in education
Show others...
2026 (English)In: Discover Artificial Intelligence, E-ISSN 2731-0809, Vol. 6, article id 166Article in journal (Refereed) Published
Abstract [en]

The assessment of open-ended written work is of vital importance to the student learning experience. Conventional essay grading methods heavily depend on expert manual assessment, making them susceptible to errors due to fatigue, bias, and subjectivity. To address this, recent research has introduced AI-based Automated Essay Scoring (AES) systems. While most studies have concentrated on predicting scores, only a few have integrated AES systems with the well-known Large Language Models (LLMs). This study explores the application of LLMs, including GPT and Gemini for AES. The proposed approach was evaluated on two benchmark datasets, namely “Hewlett Foundation: Automated Essay Scoring (ASAP–AES)” and “Learning Agency Lab–Automated Essay Scoring 2.0 (LA–AES)”. The proposed method achieved promising results in AES, demonstrating effectiveness on both the benchmark datasets. Statistical analysis revealed that Gemini outperformed GPT, achieving an average Quadratic Weighted Kappa (QWK) score of 0.45 on the ASAP–AES and 0.43 on the LA–AES. To assess the generalizability and objectivity of the proposed approach, real-world data was collected from an O-Level classroom at Sukkur IBA Community College, Pakistan. Multiple human evaluators participated in the study to examine potential biases in human assessment. The findings indicate that LLM-based scoring demonstrates improved objectivity and reduced bias compared to human assessors.

Place, publisher, year, edition, pages
Springer Nature, 2026
Keywords
LLMs, Education, Assessment, Transformers, Writing evaluation, Automated essay scoring
National Category
Computer and Information Sciences
Research subject
Computer and Information Sciences Computer Science, Computer Science
Identifiers
urn:nbn:se:lnu:diva-145293 (URN)10.1007/s44163-026-01002-y (DOI)2-s2.0-105030708166 (Scopus ID)
Available from: 2026-02-27 Created: 2026-02-27 Last updated: 2026-04-22Bibliographically approved
Taj, S., Daudpota, S. M., Imran, A. S. & Kastrati, Z. (2025). Aspect-based sentiment analysis for software requirements elicitation using fine-tuned Bidirectional Encoder Representations from Transformers and Explainable Artificial Intelligence. Engineering applications of artificial intelligence, 151, Article ID 110632.
Open this publication in new window or tab >>Aspect-based sentiment analysis for software requirements elicitation using fine-tuned Bidirectional Encoder Representations from Transformers and Explainable Artificial Intelligence
2025 (English)In: Engineering applications of artificial intelligence, ISSN 0952-1976, E-ISSN 1873-6769, Vol. 151, article id 110632Article in journal (Refereed) Published
Abstract [en]

Aspect-Based Sentiment Analysis (ABSA) of app reviews allows a better understanding of user preferences regarding specific product features and helps the development team elicit requirements effectively. The existing literature faces challenges such as limited focus on the automation of Requirement Elicitation (RE), insufficient task-specific fine-tuning of models such as Bidirectional Encoder Representations from Transformers (BERT), and lack of interpretability owing to the black-box nature of these models. Therefore, our work makes the following significant contributions to address these challenges: (1) development and evaluation of a robust method based on ABSA for the automation of the RE process; (2) optimization of ABSA using BERT fine-tuning for enhanced performance, which includes conducting a comprehensive ablation study to obtain the best hyperparameters that guarantee the best model performance and robustness; and (3) integration of Explainable Artificial Intelligence (XAI) techniques for enhanced BERT model interpretability. Our work was evaluated on the ABSA Warehouse of Apps REviews (AWARE) dataset, a specifically tailored dataset for the RE process. Our study outperformed baseline models such as the Support Vector Machine (SVM), Convolutional Neural Network (CNN), and BERT, and achieved an average F1-Score of 0.83 for the Aspect Category Detection (ACD) task and 0.94 for the Aspect Category Polarity (ACP) task. In addition, we employed XAI using Locally Interpretable Model-Agnostic Explanations (LIME) to explain the BERT model prediction results, which aids in the improved visualization and interpretability of the app review analysis for the automated RE process.

Place, publisher, year, edition, pages
Elsevier, 2025
Keywords
Software requirement elicitation, Sentiment analysis, Aspect Category Detection, Aspect Category Polarity, App reviews, Fine-tuned Bidirectional Encoder Representations from Transformers, Explainable Artificial Intelligence
National Category
Natural Language Processing
Research subject
Computer and Information Sciences Computer Science, Information Systems
Identifiers
urn:nbn:se:lnu:diva-137445 (URN)10.1016/j.engappai.2025.110632 (DOI)001459509700001 ()2-s2.0-105001036531 (Scopus ID)
Available from: 2025-03-27 Created: 2025-03-27 Last updated: 2026-04-16Bibliographically approved
Mawaldi, M. H., Kastrati, Z. & Gustafsson, A. (2025). AWACopilot: A Secure On-Premise Large Language Model-Based Solution for Enhanced Patent Drafting. In: PatentSemTech 2025: Patent Text Mining and Semantic Technologies 2025: . Paper presented at PatentSemTech 2025: Patent Text Mining and Semantic Technologies 2025. CEUR-WS, 4062
Open this publication in new window or tab >>AWACopilot: A Secure On-Premise Large Language Model-Based Solution for Enhanced Patent Drafting
2025 (English)In: PatentSemTech 2025: Patent Text Mining and Semantic Technologies 2025, CEUR-WS , 2025, Vol. 4062Conference paper, Published paper (Refereed)
Abstract [en]

Patent drafting is a complex and high-stakes process for securing intellectual property rights. During the patent prosecution phase, maintaining confidentiality is crucial, making cloud-based third-party services inadequate for patent drafting assistance due to data security concerns. This study proposes AWACopilot, a secure, on-premise solution comprising a web service that leverages open-source large language models (LLMs) to assist patent attorneys in the intricate patent application drafting process. AWACopilot generates key patent sections such as background, abstract, detailed description, etc., from human-crafted claims, addressing the data security risks posed by cloud-based AI services. Its modular architecture enables customization and adaptability to different patent tasks. Although challenges remain-including reliance on LLM capabilities and the need for rigorous content verification-this study demonstrates the potential for secure, AI-driven solutions to enhance patent drafting workflows.

Place, publisher, year, edition, pages
CEUR-WS, 2025
Series
Ceur Workshop Proceedings
Keywords
Intellectual Property, LLM, Patents Drafting, Privacy, Prompt Engineering
National Category
Information Systems
Identifiers
urn:nbn:se:lnu:diva-142923 (URN)2-s2.0-105019500317 (Scopus ID)
Conference
PatentSemTech 2025: Patent Text Mining and Semantic Technologies 2025
Available from: 2025-12-23 Created: 2025-12-23 Last updated: 2026-01-08Bibliographically approved
Ghalandarzadeh, S., Kurti, A., Unell, C., Hallborg, A., Kastrati, Z. & Sjökvist, T. (2025). Community-Based Business Models For Agricultural And Forestry Data Ecosystems: A Systematic Literature Review. Smart Agricultural Technology, 11, Article ID 100958.
Open this publication in new window or tab >>Community-Based Business Models For Agricultural And Forestry Data Ecosystems: A Systematic Literature Review
Show others...
2025 (English)In: Smart Agricultural Technology, E-ISSN 2772-3755, Vol. 11, article id 100958Article in journal (Refereed) Published
Abstract [en]

Data-driven solutions are becoming essential to modern business models, changing traditional business practices and complex value chains in multi-stakeholder and community-based sectors such as agriculture and forestry. Nevertheless, there is a lack of consolidated knowledge about the benefits and challenges that data-driven community-based business models may present in these domains. This study conducts a systematic literature review of scientific publications to identify the benefits and barriers that community-based business models for agriculture and forestry data ecosystems present. The articles included are in English and peer-reviewed and were published between 2014 to 2024. The search was conducted in Scopus, Web of Science, and IEEE Xplore, and the query resulted in 387 studies. This review has followed the PRISMA methodology, and the final number of reviewed papers was 51. Ultimately, it is found that the benefits outweigh the barriers in terms of their repetition across the literature. Significant benefits identified are interconnectedness and interactivity, resource availability, and multidirectional knowledge transfer, while the high cost of implementation and the complexity of integration and implementation of data-driven community-based business models are among the major barriers. The findings from this work can bridge the existing gap of attention to data-driven community-based business models in agriculture and forestry, help with future research work, and act as guidelines for implementing such business models.

Place, publisher, year, edition, pages
Elsevier, 2025
Keywords
Community-based business models, Agriculture and forestry data, Data ecosystems, Benefits, Barriers, Systematic literature review
National Category
Business Administration Agriculture, Forestry and Fisheries Information Systems
Research subject
Computer and Information Sciences Computer Science, Information Systems; Economy, Business Informatics; Technology (byts ev till Engineering), Forestry and Wood Technology
Identifiers
urn:nbn:se:lnu:diva-138142 (URN)10.1016/j.atech.2025.100958 (DOI)001481678700001 ()2-s2.0-105003377175 (Scopus ID)
Projects
EnTrust Next Generation of Trustworthy Agri-Data Management Grant agreement ID: 101073381
Funder
EU, Horizon Europe, 101073381
Available from: 2025-04-23 Created: 2025-04-23 Last updated: 2026-04-16Bibliographically approved
Fahad, M., Mobeen, N. E., Imran, A. S., Daudpota, S. M., Kastrati, Z., Cheikh, F. A. & Ullah, M. (2025). Deep insights into gastrointestinal health: A comprehensive analysis of GastroVision dataset using convolutional neural networks and explainable AI. Biomedical Signal Processing and Control, 102, Article ID 107260.
Open this publication in new window or tab >>Deep insights into gastrointestinal health: A comprehensive analysis of GastroVision dataset using convolutional neural networks and explainable AI
Show others...
2025 (English)In: Biomedical Signal Processing and Control, ISSN 1746-8094, E-ISSN 1746-8108, Vol. 102, article id 107260Article in journal (Refereed) Published
Abstract [en]

The gastrointestinal (GI) tract is critical in digestion and nutrient absorption, thus vital for human health. However, it is prone to diseases like cancer. Manual assessments introduce accuracy variations, consistency issues, and delays. Resources like the GastroVision dataset were introduced to advance AI in this field. Yet, it faces class imbalance issues, and baseline evaluation lacks novel methodologies, impacting accuracy. We propose a novel deep-learning model to enhance accuracy and robustness. Our approach involves averaging weights of multiple models fine-tuned with diverse hyper-parameters. In contrast to classical ensembles, our approach uses DenseNet-121 as a baseline and enables the averaging of numerous models without incurring extra inference or memory costs. Data augmentation techniques are incorporated to address class imbalance. We achieve promising results on standard performance metrics, substantially improving over baseline, notably 2.4% points in Macro Precision. Additionally, we integrate explainable AI (XAI) techniques to enhance reliability and interpretability, shedding light on the model's decision-making processes. Our study contributes to robust methodologies for imbalanced datasets, promoting model transparency and trust in predictive outcomes for clinical decision support systems.

Place, publisher, year, edition, pages
Elsevier, 2025
Keywords
GastroVision, Deep learning, Model soups, Model explainability, Generative AI, Explainable AI (XAI), Generative adversarial network (GAN), Convolutional neural network (CNN)
National Category
Biomedical Laboratory Science/Technology
Research subject
Health and Caring Sciences, Health Informatics
Identifiers
urn:nbn:se:lnu:diva-133748 (URN)10.1016/j.bspc.2024.107260 (DOI)001373625000001 ()2-s2.0-85211021231 (Scopus ID)
Available from: 2024-12-04 Created: 2024-12-04 Last updated: 2026-04-16Bibliographically approved
Fetahi, E., Susuri, A., Hamiti, M., Kastrati, Z., Canhasi, E. & Misini, A. (2025). Enhancing social media hate speech detection in low-resource languages using transformers and explainable AI. Social Network Analysis and Mining, 15, Article ID 82.
Open this publication in new window or tab >>Enhancing social media hate speech detection in low-resource languages using transformers and explainable AI
Show others...
2025 (English)In: Social Network Analysis and Mining, ISSN 1869-5450, E-ISSN 1869-5469, Vol. 15, article id 82Article in journal (Refereed) Published
Abstract [en]

Hate speech (HS) on social media exposes public discourse and community well-being. Despite its prevalence, majority of the research and technological advances have focused on rich-resourced languages. In contrast, Albanian, a language marked by dialects, limited NLP tools, and scarcity of annotated data, remains underexplored. To address this gap, this paper explores the enhancement of automatic hate speech detection in Albanian social media using advanced deep neural techniques. We collected a real-life dataset comprising 20,860 Facebook comments manually annotated using a rigorous multi-annotator process. Various techniques, including machine learning, deep neural networks, and transformers, are tested on the collected dataset. Our findings show that character-level 4-gram TF-IDF consistently reinforces ML classifiers effective with informal spelling and slang. However, the fine-tuned transformer XLM-RoBERTa achieves the highest performance with an F1-score up to 86%. These results mark a new benchmark for Albanian HS detection and highlight the impact of thorough feature engineering and robust representation techniques in a low-resource context. We also carried out an error analysis using both manual inspection and explainable AI approaches such as SHAP and LIME. Key language challenges include dialectal writing, short sentences, and implicit insults. These insights highlight the importance of domain-aligned embedding resources, more extensive annotation, and fine-grained context handling for improved detection. This study establishes a new state-of-the-art result for Albanian hate speech detection, provides a valuable comparison of models and feature engineering, and underscores the potential of multilingual transformers in low-resource scenarios.

Place, publisher, year, edition, pages
Springer, 2025
Keywords
Social media, Hate speech, Transformers, Machine learning, Deep learning, Feature engineering, Explainable AI (XAI)
National Category
Computer Sciences Natural Language Processing
Research subject
Computer and Information Sciences Computer Science, Information Systems
Identifiers
urn:nbn:se:lnu:diva-141052 (URN)10.1007/s13278-025-01497-w (DOI)001541394900003 ()2-s2.0-105014110779 (Scopus ID)
Available from: 2025-08-12 Created: 2025-08-12 Last updated: 2026-04-16Bibliographically approved
Sivakumar, M., Imran, A. S., Kastrati, Z., Soylu, A. & Kastrati, M. (2025). The Future of Grading: Can LLMs Accurately Score Student Work?. In: Proceedings of the Future Technologies Conference (FTC) 2025, Volume 4: . Paper presented at Future Technologies Conference (pp. 268-289). Springer Nature
Open this publication in new window or tab >>The Future of Grading: Can LLMs Accurately Score Student Work?
Show others...
2025 (English)In: Proceedings of the Future Technologies Conference (FTC) 2025, Volume 4, Springer Nature, 2025, p. 268-289Conference paper, Published paper (Refereed)
Abstract [en]

Exploring the transformative potential of large language models (LLMs) in revolutionizing the assessment process in education is a pressing need. LLMs have the capability to automatically evaluate student submissions, significantly enhancing the educational landscape and providing a second opinion and additional support to human evaluators. Therefore, our evaluation of various LLMs aims to assess their ability to provide automated, detailed, consistent, unbiased, and efficient feedback and scoring in educational settings, instilling a sense of optimism about the future of assessment. To achieve our objective, we comprehensively evaluated various state-of-the-art LLMs, such as GPT-4, GPT-4o, and Mixtral 8x22B on diverse datasets, including ASAP SAS, ASAP AES, and a real-world BWD dataset specifically collected for this study. The experimental results on these datasets employing sound prompt engineering techniques demonstrate that LLMs possess the potential not only to automate the scoring of student submissions but also to accurately match the scores of human assessors for actual courses taught in universities. Notably, GPT-4o exhibited promising capabilities in scoring short-answer submissions. These models were particularly good in STEM-related domains for tasks with clear structure, well-defined rubrics, and minimal subjective interpretation. However, the study identifies specific challenges for certain tasks, particularly creative tasks, underscoring the need for further research in this area.

Place, publisher, year, edition, pages
Springer Nature, 2025
National Category
Natural Language Processing
Identifiers
urn:nbn:se:lnu:diva-142224 (URN)10.1007/978-3-032-07992-3_18 (DOI)2-s2.0-105021830683 (Scopus ID)9783032079923 (ISBN)
Conference
Future Technologies Conference
Available from: 2025-10-28 Created: 2025-10-28 Last updated: 2026-01-21Bibliographically approved
Kastrati, M., Imran, A. S., Hashmi, E., Kastrati, Z., Daudpota, S. M. & Biba, M. (2025). Unlocking language barriers: Assessing pre-trained large language models across multilingual tasks and unveiling the black box with Explainable Artificial Intelligence. Engineering applications of artificial intelligence, 149, Article ID 110136.
Open this publication in new window or tab >>Unlocking language barriers: Assessing pre-trained large language models across multilingual tasks and unveiling the black box with Explainable Artificial Intelligence
Show others...
2025 (English)In: Engineering applications of artificial intelligence, ISSN 0952-1976, E-ISSN 1873-6769, Vol. 149, article id 110136Article in journal (Refereed) Published
Abstract [en]

Large Language Models (LLMs) have revolutionized many industrial applications and paved the way for fostering a new research direction in many fields. Conventional Natural Language Processing (NLP) techniques, for instance, are no longer necessary for many text-based tasks, including polarity estimation, sentiment and emotion classification, and hate speech detection. However, training a language model for domain-specific tasks is hugely costly and requires high computational power, thereby restricting its true potential for standard tasks. This study, therefore, provides a comprehensive analysis of the latest pre-trained LLMs for various NLP-related applications without fine-tuning them to evaluate their effectiveness. Five language models are thus employed in this study on six distinct NLP tasks (including emotion recognition, sentiment analysis, hate speech detection, irony detection, offensiveness detection, and stance detection) for 12 languages from low- to medium- and high-resource. Generative Pre-trained Transformer 4 (GPT-4) and Gemini Pro outperform state-of-the-art models, achieving average F1 scores of 70.6% and 68.8% on the Tweet Sentiment Multilingual dataset compared to the state-of-the-art average F1 score of 66.8%. The study further interprets the findings obtained by the LLMs using Explainable Artificial Intelligence (XAI). To the best of our knowledge, it is the first time any study has employed explainability on pre-trained language models.

Place, publisher, year, edition, pages
Elsevier, 2025
Keywords
Large language models, Zero-shot classification, Explainable Artificial Intelligence, Sentiment analysis, Emotion recognition
National Category
Natural Language Processing
Research subject
Computer and Information Sciences Computer Science, Information Systems
Identifiers
urn:nbn:se:lnu:diva-137213 (URN)10.1016/j.engappai.2025.110136 (DOI)001446400200001 ()2-s2.0-86000570396 (Scopus ID)
Available from: 2025-03-12 Created: 2025-03-12 Last updated: 2026-04-16Bibliographically approved
Kastrati, Z., Fatima, S., Kurti, A., Daudpota, S. M. & Imran, A. S. (2024). Analyzing and Predicting the Helpfulness of Reviews in MOOCs Context Using Deep Learning. In: Jonathan Flearmoy (Ed.), Procedia Computer Science 246: 28th International Conference on Knowledge Based and Intelligent information and Engineering Systems (KES 2024). Paper presented at 28th International Conference on Knowledge Based and Intelligent information and Engineering Systems (KES 2024) (pp. 772-781). Elsevier, 246
Open this publication in new window or tab >>Analyzing and Predicting the Helpfulness of Reviews in MOOCs Context Using Deep Learning
Show others...
2024 (English)In: Procedia Computer Science 246: 28th International Conference on Knowledge Based and Intelligent information and Engineering Systems (KES 2024) / [ed] Jonathan Flearmoy, Elsevier, 2024, Vol. 246, p. 772-781Conference paper, Published paper (Refereed)
Abstract [en]

Students' feedback is an essential part of the teaching-learning process and serves as an effective instrument for continuous improvement in educational environments. The insights gathered from students' experiences and perceptions expressed in reviews provide instructors with a valuable resource to enhance their teaching methods, instructional design, and overall classroom dynamics. However, students' reviews are often unclear, contradictory, and conflicting with each other, making their interpretation and use challenging. Therefore, this study proposes a novel deep learning-based approach that helps course designers and instructors effectively identify constructive and useful reviews. The approach leverages the integration of several attributes, including textual review, student satisfaction, meta-data of the course, and review-derived information such as sentiment, readability, and review depth. The approach is tested on a real-life dataset comprising 38,717 reviews gathered from the Coursera learning platform for the purpose of this study. The experimental results, with an F1-score of 0.91, suggest that the approach can be an effective tool for educators and instructional designers to identify helpful student reviews.

Place, publisher, year, edition, pages
Elsevier, 2024
Series
Procedia Computer Science, E-ISSN 1877-0509 ; 246
Keywords
Review helpfulness, MOOCs, Deep learning, Student satisfaction, Student's feedback, Review depth, Sentiment, Meta-data
National Category
Information Systems
Research subject
Computer and Information Sciences Computer Science, Computer Science
Identifiers
urn:nbn:se:lnu:diva-133565 (URN)10.1016/j.procs.2024.09.496 (DOI)2-s2.0-85213356696 (Scopus ID)
Conference
28th International Conference on Knowledge Based and Intelligent information and Engineering Systems (KES 2024)
Available from: 2024-11-29 Created: 2024-11-29 Last updated: 2026-04-07Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-0199-2377

Search in DiVA

Show all publications