Expert Opinion Form on the Use of Artificial Intelligence in Assessment and Evaluation in Moodle


Bulletin of the Technical Committee on Learning Technology (ISSN: 2306-0212)
Volume 26, Number 1, 33–38 (2026) Download PDF
Received June 30, 2026
Accepted August 2, 2026
Published online August 9, 2026
This work is under the Creative Commons CC BY-NC-ND 3.0 license. For more information, see Creative Commons License .


Authors:Zehra Altinay1, Fahriye Altinay1, Zöhre Serttaş1, Gökmen Dagli2, and Sezer Kanbul1

1: Near East University, Nicosia, North Cyprus, Mersin 10 Turkiye; 2: University of Kyrenia, Kyrenia, North Cyprus, Mersin 10 Turkiye


Abstract:

Given the rapid technological advancement, this qualitative study aims to explore expert opinions on the integration, affordances and challenges of artificial intelligence (AI) in the assessment and evaluation functions of Moodle. Data collection was carried out by means of semi-structured opinion forms from an international sample of 20 senior experts in educational technology. Data were analysed using inductive thematic analysis. The results suggest a dualistic landscape with affirmative dimensions that show how AI boosts operational efficiency, scales personalized feedback loops, and streamlines dynamic item generation and cautious dimensions that reveal critical bottlenecks, such as domain-specific unreliability (STEM vs humanities), semantic gaming, and potential algorithmic bias. This study demonstrates that such epistemic and ethical limits suggest that the prospect of grading by full automation is premature. Therefore, we recommend a calibrated “human-in-the-loop” hybrid paradigm where AI serves as an administrative and formative catalyst while human educators retain the ultimate evaluative authority to ensure equity.

Keywords: Artificial Intelligence; Moodle LMS; Educational Assessment.

I. INTRODUCTION

As we strive to keep pace with the rapid advancements in technology in our modernising world, the technologies used in education are also evolving. The Moodle platform, in particular, is widely used within the educational community as it offers a range of features that facilitate the uploading of lecture notes, the delivery of interactive live lessons, and the submission and assessment of assignments. Thanks to its features, it also provides tools that support assessment and evaluation (Grévisse, Braun, & Batista da Costa, 2025). Thanks to advancing technologies, the use of Artificial Intelligence has also come to the fore on this platform. It is argued that Artificial Intelligence—a topic that has been at the forefront of research in recent years—offers the potential for personalised learning; when combined with Moodle’s functional features for identifying students’ weaknesses in lessons and assessing their performance, this is expected to enhance the quality of assessment and evaluation (Kaleci , 2025; Heil, Ifenthaler, & Cooper, 2025).

On the other hand, the limitations of artificial intelligence in education should not be overlooked. Research in the literature highlights various debates regarding the use of artificial intelligence in education. Despite all these promising developments, integrating artificial intelligence with Moodle-based assessment still poses challenges. The use of AI can increase efficiency, simplify the task of grading which is the responsibility of educators and speed up the process of providing feedback; however, concerns such as the reliability of grades and potential algorithmic biases are slowing down the adoption of these systems. In short, a review of existing studies suggests that while AI is viewed as a supportive tool, responsibility for critical assessment stages must remain with humans. In this context, the “Human in the Loop” (HITL) approach has emerged as a strategy aimed at combining human judgment with AI for assessment in education. The scarcity of research in the literature discussing the applicability of this approach within Moodle-based assessment systems serves as the motivation for this study (Stokkink, 2025).

This study, titled “Expert Opinion Form on the Use of Artificial Intelligence in Assessment and Evaluation in Moodle”, aims to identify expert opinions regarding the use of Artificial Intelligence for assessment functions on the Moodle platform. The research is expected to contribute to the literature on the ethical and effective use of Artificial Intelligence in the assessment and evaluation functions of Moodle.

II. METHODOLOGY

A. Research Design

The study uses a qualitative descriptive research design to collect, analyze and organize the subtle expert opinions on the integration, effectiveness and challenges of Artificial Intelligence (AI) in the assessment and evaluation processes of the Moodle Learning Management System (LMS). This investigation is particularly suited to qualitative description, which enables researchers to gather direct responses to questions of immediate interest to educational practitioners and policymakers, without too much theoretical abstraction and while preserving the operational richness of experts’ insights.

B. Participants and Sampling

We employed purposive sampling strategy to ensure the high-level expertise and data richness. The sample included 20 experts holding academic positions such as professors, senior lecturers, and senior researchers. The participants were affiliated with universities in the Turkish Republic of Northern Cyprus (TRNC), Turkey, China, the United Kingdom, the United States, and Romania. Their areas of expertise included Educational Technology, Computer Science, Artificial Intelligence in Education, Measurement and Evaluation, and Educational Psychology. The participants have between 5 and 25 years of experience in measurement and evaluation. Theoretical saturation was reached and data collection ceased with a total sample of N = 20 experts from different international institutions of higher education.

TABLE I EXPERT INFORMATION LIST

Participant CodeAcademic TitleField of ExpertiseInstitutional LocationAssessment Experience (Years)
E1Associate ProfessorEducational TechnologyChina5
E2Professor, Dr.Instructional Design & AITurkey23
E3Full ProfessorComputer Science & LMS ArchitectureCyprus18
E4Assistant ProfessorLearning AnalyticsUK7
E5Senior LecturerAssessment & EvaluationChina12
E6Professor, Dr.Distance EducationCyprus20
E7Associate ProfessorNatural Language Processing in Ed.Cyprus9
E8Lecturer, PhDInstructional TechnologyUK6
E9Full ProfessorCurriculum and InstructionChina25
E10Associate ProfessorDigital AssessmentCyprus11
E11Professor, Dr.Informatics & Data ScienceTurkey15
E12Assistant ProfessorOnline PedagogiesCyprus8
E13Senior ResearcherCognitive Technologies in Ed.Cyprus14
E14Full ProfessorEducational PsychologyUSA22
E15Associate ProfessorAssessment MetricsRomania10
E16Lecturer, PhDE-Learning SystemsCyprus7
E17Professor, Dr.Artificial Intelligence in Ed. (AIEd)Cyprus16
E18Associate ProfessorInstructional DesignTurkiye13
E19Assistant ProfessorEducational Software Dev.Turkiye6
E20Full ProfessorOpen and Distance LearningUK21
C. Data Collection Instrument and Procedure

The data was collected through a semi-structured open-ended Expert Opinion Form with 10 broad items, which investigated the systemic, pedagogical, technical and ethical aspects of AI-based evaluation in Moodle. The open-ended questions were developed in response to a review of emergent literature on GenAI tools and educational measurement. Prior to administration, the instrument was checked for content validity by two senior independent experts in educational measurement. Data were collected digitally during the period February to June 2026. All participants gave informed consent before responding.

D. Data Analysis

The 20 expert forms were analysed using inductive thematic analysis, following the multi-step framework implemented by Braun and Clarke. First, all experts’ responses were read and examined in detail. Next, open coding was performed to identify meaningful concepts and recurring patterns found in many of the responses. Then, similar appearing codes were grouped into larger categories. To increase the reliability of the analyses, a subset of the data was coded by two researchers independently and their results were compared. Discrepancies were discussed until consensus was reached and the coding framework was systematically applied to all data.

The analytical pipeline was:

  • Repeated active reading to familiarize with transcripts.
  • Development of first open codes to expose overt meanings and covert patterns.
  • Categorizing codes into broader, separate categories based on structural resemblances.
  • Sort and re-sort to identify higher order themes that are superordinate and directly relate to Positive/Affirmative (Opportunities) and Negative/Cautious (Challenges/Limitations) outlooks.
  • Theme names finalised and representative verbatims extracted.

To verify the credibility of the data, inter-coder reliability was established by two researchers independently coding 30% of the dataset. Disagreements were resolved through consensus after discussion and a Miles and Huberman reliability coefficient of 91.5% was obtained.

III. FINDINGS AND DISCUSSION

The systematic thematic analysis of the responses of the 20 experts produced 4 main themes and 12 sub-categories. The results are presented as Affirmative Dimensions (Opportunities and Pedagogical Affordances) and Cautious Dimensions (Challenges, Systemic Bottlenecks and Ethical Constraints) to mirror the dualistic nature of AI deployment in Moodle’s evaluation infrastructure.

TABLE II. Expert Information List

Macro OutlookMain ThemeSub-Categories / Core CodesFrequency (f)*
Affirmative Dimensions (Positive)Theme 1: Pedagogical and Operational Affordances1.1. Operational Efficiency and Automation of Routine Tasksf = 19
1.2. Scalable Personalized Feedback Ecosystemsf = 18
1.3. Enhancement of Objectivity and Standardizationf = 14
Theme 2: Exam Engineering and Predictive Analytics2.1. Dynamic Item Generation and Parallel Assessment Designf = 16
2.2. Psychometric Modeling and Historical Data Trackingf = 11
Cautious Dimensions (Negative)Theme 3: Epistemic and Technical Limitations3.1. Domain-Specific Reliability Degradation (STEM vs. Humanities)f = 15
3.2. Technical Interoperability and Upfront Integration Costsf = 13
3.3. Vulnerability to Phrasing, Gaming, and Semantic Inaccuracyf = 12
Theme 4: Ethical Safeguards and Academic Integrity Disruption4.1. Algorithmic Bias Perpetuation and Equity Concernsf = 14
4.2. Evolutionary Shifts in Academic Dishonesty (GenAI Exploit)f = 17
4.3. Imperative of “Human-in-the-Loop” (Expert Oversight)f = 20

*Note: Frequencies refer to the number of experts (N=20) that explicitly mentioned codes that contributed to the respective sub-categories.

A. Theme 1: Pedagogical and Operational Affordances (Positive Outlook)

This theme emphasizes the structural and pedagogical value that AI has for the Moodle ecosystem. Experts overwhelmingly agreed that AI tools democratize the delivery of feedback, reducing the cognitive and time burden on instructors.

Fig. 1. Theme 1: Pedagogical & Operational Affordances

1)    Category 1.1: Operational Efficiency and Automation of Routine Tasks

Our data support that the AI is most appropriate for the tasks with structured and rule based grading mechanisms. This saves a huge amount of time dealing with large groups of students.

“AI can handle well structured tasks (like multiple-choice, code output, numeric responses)… overall AI-assisted assessment is promising for speed and consistency.” (Expert 1)

“The more the instructor gets used to it, the more automated, less time-consuming it becomes… it gives fast results.” (Expert 2)

“For Moodle general education classes with large enrollment, automatic screening of baseline answers saves dozens of hours of repetitive labor. (Expert 12)

2) Category 1.2: Scalable Personalized Feedback Ecosystems

The experts say that one of the main benefits of AI is that it can provide detailed, real-time feedback – something teachers cannot do for each student due to time constraints.

“It gives very specific feedback that an instructor cannot afford to give to every student.” (Expert 2)

“AI means instant, personalized feedback, especially for procedural mistakes… These speeds up the feedback loop and reduces student anxiety waiting for feedback.” (Expert 1)

“Students in online settings need immediate reinforcement. Waiting two weeks for a human tutor to grade an essay kills learning. AI fills this gap.”  (Expert 6)

3) Category 1.3: Enhancement of Objectivity and Standardization

Experts pointed out that “AI minimizes human subjectivity, grading fatigue and subconscious biases that often plague manual assessment processes.”

“I think objectivity could go up with the AI tools.” (Expert 2).

“AI can increase objectivity by applying the same criteria to all students, without grader fatigue or inconsistency.” (Expert 1).

B. Theme 2: Exam Engineering and Predictive Analytics (Positive Outlook)

This theme reports on the contributions of AI in the pre-assessment phase, in particular on content generation, item bank optimization and structural alignment within Moodle.

1)    Category 2.1: Dynamic Item Generation and Parallel Assessment Design

Generative AI in Moodle enables quick generation of various evaluation metrics that are contextually and structurally consistent.

“AI helps by developing question banks based on learning objectives… creating alternate versions of exams to minimize cheating, and proposing scenario-based or case-based questions. (Expert 1)

“By planning the exam beforehand and directing the AI tools according to expectations… the results could be very helpful.” (Expert 2).

The ability of AI to generate instant contextual variations of a single problem schema greatly diminishes the utility of student answer leaks. (Expert 17).

2)    Category 2.2: Psychometric Modeling and Historical Data Tracking

AI can analyze historical student data in Moodle to help improve exam design in the future, Amdocs said.

“An important affordance is the estimation of difficulty and discrimination indices from historical data.” (Expert 1).

The AI can not only write questions, but it can also predict item performance based on past logs in the moodle database. (Expert 11)

C. Theme 3: Epistemic and Technical Limitations (Negative Outlook)

The experts have expressed serious concerns about the technical realization and epistemological limitations of the current AI models integrated within Moodle, while acknowledging the obvious advantages.

Fig. 2.  Theme 3: Epistemic & Technical Limitations

1)    Category 3.1: Domain-Specific Reliability Degradation

The success of automated evaluation is highly polarized across the academic disciplines. It is very accurate on algorithmic or mathematical things but less accurate on interpretive tasks.

Reliability also differs across subject domains – more reliable in STEM calculation tasks, less so in humanities interpretation. (Expert 1).

Its accuracy is moderate and essays and complex open-ended questions need human moderation… because they don’t do well in assessment of the reasoning process, originality and coherence of an argument (Expert 1).

“AI models are good at evaluating deterministic code execution or mathematical proofs in Moodle, but not at grasping deep philosophical nuances in humanities portfolios. (Expert 3)

2)    Category 3.2: Technical Interoperability and Initial Integration Cost

The shift to AI-enabled assessment across the system will include workflow changes, manual preparation bottlenecks, and financial parameters that are often underestimated.

If the exam papers are in print (not in electronic forms) then it could be extra time to convert them into electronic forms. And you might need to spend a little time initially preparing the prompts and the rubrics. (Expert 2)

“Technical issues: stability of integration with Moodle, handling of multimedia response…and API costs for cloud-based AI services. (Expert 1)

“It takes a lot of faculty effort up front to build out robust system infrastructures and validate the initial prompts. (Expert 15)

3)    Category 3.3: Vulnerability to Phrasing, Gaming and Lack of Semantic Accuracy

Current NLP implementations for LMS frameworks often rely on superficial textual cues rather than on actual conceptual understanding.

“Reliability is compromised since it is sensitive to phrasing, length and surface level patterns” (e.g. keyword matching without understanding). (Expert 1)

“Students can game the grading model, filling essays with certain pedagogical buzz words that trick the AI into giving them a better mark. (Expert 19)

C. Theme 4: Ethical safeguards and disruption of academic integrity (Negative outlook)

The theme outlines the basic socio-ethical problems, weaknesses of institutional policy and the non-negotiable requirement of human governance over the automated systems.

1)    Category 1: Algorithmic bias Equity and Perpetuation Concerns

Since AI models are trained on historical data, they may inadvertently formalize and accelerate existing socio-linguistic biases in Moodle’s grading pathways.

“Fairness is not baked in – if the training data or the rubric has biases (e.g. penalising non-native English style), the AI will reinforce those biases. (Expert 1).

“AI grading that’s standardized could miss important accessibility issues for students with disabilities or spotty internet connections.” (Expert 5).

2)    Category 4.2: Evolutionary Shifts in Academic Dishonesty

We are implementing AI-assisted assessment in an adversarial environment where generative AI is both a cheat and an enforcer.

“Generative AI opens up new forms of academic dishonesty – students using AI to write essays or solve problems without understanding. “Assessment design has to change. (Expert 1)

“Today, it is difficult to get AI out of our daily lives. Therefore, we need new ways to adapt our assessment and style of homework assignment.” (Expert 2)

3)    Category 4.3: The Imperative of “Human-in-the-Loop” (Expert Oversight)

A consensus emerged among all 20 experts: under no circumstances should AI be granted ultimate, unmonitored decision-making authority over high-stakes summative evaluations.

Fig. 3. Human in the loop safeguard

Results should be interpreted by an expert, and final decisions should be made by the expert. Otherwise, there might be some unexpected conflicts in case of uncontrolled usage. (Expert 2).

“I suggest that automated grading be used only for initial grading or formative feedback and that instructors review borderline or exceptional cases. Human review is necessary.” (Expert 1).

“AI should be a supportive system, not a replacement for pedagogical empathy and expert judgment. (Expert 14).

Overall, while the findings suggest that human judgment is necessary ensure the reliability of AI-supported assessment in Moodle, they also indicate that artificial intelligence is perceived as a tool that enhances the efficiency and quality of assessment. Additionally, the findings suggest that artificial intelligence should support educators rather than replace them.

In future studies, the technical feasibility of AI-supported processes such as dynamic question generation should also be examined, and the proposed framework should be evaluated using indicators such as consistency and reliability.

IV. CONCLUSION & IMPLICATIONS FOR MOODLE ARCHITECTURE

The synthesis of the views of 20 international experts shows that the integration of artificial intelligence into the Moodle assessment and evaluation framework provides significant opportunities in terms of improving the operational efficiency, increasing the speed of feedback and ensuring structural standardization. But these two technologies need to be well integrated because of issues like domain-specific reliability issues, algorithmic biases and semantic security vulnerabilities. This result is consistent with the literature on the use of artificial intelligence in education. Reliability and validity of AI-based assessment systems may vary, in particular, depending on the characteristics of the domain where they are used. Artificial intelligence has promising potential for educational assessment and evaluation, but Ho (2024) warns that it also has the potential to generate risks of biased scores, construct underrepresentation, and differential impacts that could develop over time. Thus, the author emphasizes the need for continuous bias monitoring and quality assurance processes in AI-based assessment systems.

To ensure successful integration of artificial intelligence in assessment systems, institutions should not use fully automated AI-powered processes in high stakes assessments. The preferred approach is instead a hybrid model, where human actors play an active role in the assessment process all along. In this sense, artificial intelligence can be used to perform routine tasks and allow rapid formative assessment cycles. However, the final responsibility for diagnosis, judgement and final evaluation is up to human experts.

The experts’ recommendation for a hybrid approach aligns well with the growing body of literature on Human-in-the-Loop (HITL) artificial intelligence. There is a policy gap between the complete ban and the use of AI without any restriction in academic assessment, which can be well managed by human-supervised models (Khan et al. 2026). The authors argue that supervision by humans is necessary to maintain academic standards and the integrity of assessment practices. Similarly, Morgan et al. (2023) recommend that oversight for high-risk AI applications should not be left to individual decision makers but should be built into organizational mechanisms led by teams of experts. This view supports the idea that human experts should have the final word, particularly when dealing with high-stakes assessment. The results of the present study support and reinforce the conclusions reported in the literature.

In addition, future iterations of the Moodle plugins should aim to develop transparent and explainable artificial intelligence (AI) architectures, robust rubric validation processes, and comprehensive data governance frameworks to uphold student privacy and equity. The experts’ focus on explainable AI fits into the wider academic conversation around establishing reliable artificial intelligence systems in education. Baker and Hawn (2022) call for greater transparency and auditability in the design of educational algorithms to detect, analyze, and mitigate algorithmic biases. Moreover, a large body of research on Human-in-the-Loop (HITL) systems has shown that good oversight by humans does not necessarily mean that human actors are involved in decision-making. Instead, it relies on establishing system architectures that are by design explainable, auditable and able to provide interpretable justifications for algorithmic outputs and recommendations. These features are viewed as vital prerequisites to ensure accountability, develop stakeholder trust, and support the responsible integration of artificial intelligence (AI) tools in educational assessment and evaluation practices. The findings are restricted to the experts’ valuable insights. They are not validated by use in an actual classroom situation. Future research should test the proposed approach in the real Moodle settings with research designs that include different methods and test its actual effectiveness.

Reference

[1] C. Grévisse, C. Braun, and J. Batista da Costa, “AutoTag & TagMap: LLM-Powered Moodle Plugins for Pedagogical Alignment Checks,” SN Comput. Sci., vol. 6, Art. no. 813, 2025, doi: 10.1007/s42979-025-04348-9.
[2] J. Heil, D. Ifenthaler, M. Cooper, M. L. Mascia, and R. Conti, “Students’ perceived impact of GenAI tools on learning and assessment in higher education: The role of individual AI competence,” Smart Learn. Environ., vol. 12, Art. no. 37, 2025, doi: 10.1186/s40561-025-00395-0.
[3] D. Kaleci, “Integration and application of artificial intelligence tools in the Moodle platform: A theoretical exploration,” J. Educ. Technol. Online Learn., vol. 8, no. 1, pp. 100–111, Jan. 2025, doi: 10.31681/jetol.1595079.
[4] P. Stokkink, “The Impact of AI on Educational Assessment: A Framework for Constructive Alignment,” arXiv preprint arXiv:2506.23815, Jun. 2025, doi: 10.48550/arXiv.2506.23815.
[5] A. D. Ho, “Artificial Intelligence and Educational Measurement: Opportunities and Threats,” J. Educ. Behav. Stat., vol. 49, no. 5, pp. 715–722, Oct. 2024, doi: 10.3102/10769986241248771.
[6] B. Arslan, B. Lehman, C. Tenison, J. R. Sparks, A. A. López, L. Gu, and D. Zapata-Rivera, “Opportunities and challenges of using generative AI to personalize educational assessment,” Front. Artif. Intell., vol. 7, Art. no. 1460651, Oct. 2024, doi: 10.3389/frai.2024.1460651.
[7] J. A. Idowu, A. S. Koshiyama, and P. Treleaven, “Investigating algorithmic bias in student progress monitoring,” Comput. Educ.: Artif. Intell., vol. 7, Art. no. 100267, Dec. 2024, doi: 10.1016/j.caeai.2024.100267.
[8] W. Khan, L. Topham, N. Jones, P. Atherton, R. Al-Shabandar, H. Kolivand, I. Khan, A. Alatrany, and A. Hussain, “Auto-assessment of assessment: A human-in-the-loop AI framework addressing policy gaps in academic assessment,” PLOS ONE, vol. 21, no. 4, Art. no. e0346815, Apr. 2026, doi: 10.1371/journal.pone.0346815.
[9] C. Grévisse, C. Braun, and J. Batista da Costa, “AutoTag & TagMap: LLM-Powered Moodle Plugins for Pedagogical Alignment Checks,” SN Comput. Sci., vol. 6, Art. no. 813, 2025, doi: 10.1007/s42979-025-04348-9.
[10] D. Morgan, Y. Hashem, J. Francis, S. Esnaashari, V. J. Straub, and J. Bright, “‘Team-in-the-loop’: Ostrom’s IAD framework ‘rules in use’ to map and measure contextual impacts of AI,” arXiv preprint arXiv:2303.14007, Mar. 2023, doi: 10.48550/arXiv.2303.14007.
[11] R. S. Baker and A. Hawn, “Algorithmic Bias in Education,” Int. J. Artif. Intell. Educ., vol. 32, no. 4, pp. 1052–1092, Dec. 2022, doi: 10.1007/s40593-021-00285-9.
[12] K. Lazaros, A. G. Vrahatis, and S. Kotsiantis, “Human-in-the-Loop Artificial Intelligence: A Systematic Review of Concepts, Methods, and Applications,” Entropy, vol. 28, no. 4, Art. no. 377, Mar. 2026, doi: 10.3390/e28040377.


 Authors


Zehra Altinay

received the B.A., M.A., and Ph.D. degrees in educational sciences from Near East University, Nicosia, TRNC, Mersin 10, Türkiye.

She is currently a professor whose research interests include educational leadership, digital transformation, artificial intelligence in education, instructional technologies, and quality assurance in higher education.


Fahriye Altinay

she received the B.A., M.A., and Ph.D. degrees in educational sciences from Near East University, Nicosia, TRNC, Mersin 10, Türkiye. She is currently a professor whose research interests include educational technologies, online learning, digital transformation, inclusive education, and artificial intelligence in education


Zöhre Serttaş

She is a Ph.D. Student with the Digital Education and Information Technology Unit, Artificial Intelligence and Informatics Faculty, and the Research Center for AI and IoT, Near East University, Nicosia, TRNC, Mersin 10, Türkiye. Her research interests include educational technologies, artificial intelligence in education, digital learning, distance education, and technology integration.


Gökmen Dağlı

He received the B.A., M.A., and Ph.D. degrees in Educational Administration. He is currently a Professor with the Faculty of Education, University of Kyrenia, Kyrenia, TRNC, Mersin 10, Türkiye. His research interests include educational leadership, educational administration, school management, quality management in education, organizational behavior, and digital transformation in education.


Sezer Kanbul

He received the B.S., M.S., and Ph.D. degrees in Computer Education and Instructional Technology from Near East University, Nicosia, TRNC, Mersin 10, Türkiye. He is currently a professor whose research interests include educational technologies, artificial intelligence in education, digital transformation, distance learning, and technology-enhanced learning. He has published in SSCI-, Scopus-, and ESCI-indexed journals and has supervised numerous graduate theses in these fields