Evaluation of the STACK Quiz Analytics Moodle plugin — the research plan for a small pilot study with lecturers who use STACK, evaluating the plugin's usability, usefulness, technical performance, and potential educational value.
1. Background and Rationale
STACK generates rich data about students' interactions with automatically assessed mathematics questions. Moodle provides access to much of this information through quiz reports, statistics, attempt data, and question-level records. However, lecturers who wish to examine patterns across multiple quizzes or investigate student responses in greater detail often need to move between several Moodle reports, download data, or carry out additional analysis manually.
The STACK Quiz Analytics Moodle plugin was developed to simplify this process by bringing together selected course-level and question-level analytics within Moodle. The plugin is intended to reduce routine analytical work and help lecturers move more quickly from raw STACK data to interpretation and action.
The current version provides two main levels of analysis. Course-wide analytics are intended to help lecturers identify quizzes or questions that may require further attention, while Question Analytics allows lecturers to investigate how students responded to individual questions, including common incorrect responses and other response patterns. Detailed student-level information remains available through Moodle's existing reports rather than being reproduced unnecessarily within the plugin.
The next stage of development is a small pilot study with lecturers who use STACK. The purpose of the pilot is not only to determine whether the plugin functions technically, but also whether lecturers find the information understandable, useful, trustworthy, and actionable in realistic teaching contexts.
2. Aim and Research Questions
The overall aim of the pilot is to evaluate the usability, usefulness, technical performance, and potential educational value of the STACK Quiz Analytics plugin when used by lecturers with real or representative STACK course data.
The pilot will address the following preliminary research questions:
RQ1. Usability and usability of analytics
To what extent can lecturers access, navigate, and interpret the analytics provided by the plugin without substantial additional support?
RQ2. Efficiency and workload
To what extent does the plugin reduce the time and manual effort required to examine STACK quiz data compared with lecturers' usual approaches?
RQ3. Actionability
To what extent does the plugin help lecturers identify meaningful patterns in quizzes, questions, or student responses that could inform changes to assessment materials, feedback, teaching practice, or further research?
RQ4. Technical reliability and performance
How accurately and efficiently does the plugin operate across different course sizes, quiz structures, Moodle installations, and STACK datasets?
These questions are intentionally broad for the initial pilot. They can be refined after the first round of testing.
3. Participants and Testing Context
The initial pilot will involve a small purposive sample of lecturers and colleagues with experience using STACK in Moodle.
Where possible, participants will be drawn from different institutions and teaching contexts so that the plugin can be tested against variation in:
- course size;
- number of quizzes and questions;
- number of attempts;
- Moodle and STACK versions;
- undergraduate and other teaching contexts;
- lecturers with different levels of experience using STACK.
A first pilot group of approximately 5–10 lecturers would be sufficient for identifying major usability, interpretation, and technical issues. A larger second phase could then be undertaken after the plugin has been refined.
Participants may test the plugin using either their own existing STACK course data, where appropriate permissions are in place, or a prepared demonstration/test course.
The experimental Model Analytics component will not form part of the initial pilot. It will remain disabled while its design, validation, and data-processing requirements are considered separately.
4. Pilot Procedure
Each participant will first receive a short installation/access guide and a brief video demonstrating the main workflow of the plugin.
Lecturers will then be asked to use the plugin with a selected STACK course and complete a small number of common analytical tasks.
Suggested tasks are:
- Identify a quiz that appears to require further attention.
- Identify one question that may warrant closer investigation.
- Use Question Analytics to examine how students responded to that question.
- Identify one or more common response patterns or difficulties.
- Use the available links to inspect the corresponding Moodle report where more detailed information is needed.
- Describe whether the information would lead them to change, investigate, or reconsider anything in the question, feedback, assessment design, or teaching.
Participants will be encouraged to use the interface as naturally as possible rather than being given step-by-step instructions for every feature. This will allow the study to identify which parts of the interface are self-explanatory and where additional documentation or interface changes are needed.
For a subset of participants, the session may be observed or conducted as a short think-aloud exercise. Researchers will note where participants hesitate, misunderstand a metric, require assistance, or encounter technical problems.
After testing, participants will complete the short STACK Quiz Analytics Lecturer Feedback Survey, expected to take approximately 5–7 minutes. Where useful, selected participants may also take part in a short follow-up interview.
5. Data to Be Collected
The pilot will combine technical, behavioral, and self-reported evidence.
Lecturer feedback
The post-use survey will collect information about:
- ease of access and navigation;
- clarity of analytics and visualisations;
- usefulness of the information presented;
- perceived reduction in manual analytical work;
- perceived time savings;
- ability to identify quizzes or questions requiring attention;
- ability to understand student response patterns;
- confidence in the accuracy and trustworthiness of the analytics;
- intention to use the plugin in future;
- most useful features;
- missing features;
- problems encountered;
- suggestions for improvement.
Open-ended questions will also ask lecturers whether the plugin revealed anything they had not previously noticed and whether the results could lead to a change in teaching, assessment, question design, or further investigation.
Task-based observations
Where possible, the research team will record:
- whether each assigned task was completed;
- approximate time required;
- points where assistance was needed;
- features that were overlooked;
- analytics that were misunderstood;
- navigation difficulties;
- comments made during use.
Technical information
Technical characteristics of the test environment should be recorded separately rather than relying only on lecturer recall.
These may include:
- Moodle version;
- STACK version;
- number of enrolled students;
- number of students with attempts;
- number of quizzes;
- number of questions;
- number of quiz attempts;
- page/analysis loading times;
- timeouts or failed computations;
- discrepancies between plugin values and corresponding Moodle reports.
Particular attention will be paid to performance with larger datasets because computation time has already been identified as an area requiring further testing.
6. Data Analysis
Quantitative survey items will initially be analysed descriptively using frequencies, medians or means, and response distributions.
Results may be grouped under four principal evaluation dimensions:
Usability
- navigation;
- clarity;
- ease of finding information.
Efficiency
- perceived reduction in manual analysis;
- perceived time saving.
Actionability
- identification of problematic quizzes/questions;
- understanding of student response patterns;
- reported changes or investigations prompted by the analytics.
Trust and adoption
- perceived accuracy;
- willingness to use the plugin;
- overall usefulness.
Open-ended survey and interview responses will be reviewed thematically to identify recurring strengths, weaknesses, feature requests, interpretation problems, and examples of actionable insights.
Technical measurements will be analysed separately to identify performance bottlenecks, compatibility issues, and discrepancies with Moodle's existing reports.
The first pilot will primarily be formative: findings will be used to improve the plugin and refine the design of a larger evaluation.
7. Ethics and Data Protection
Before formal research data collection begins, the project team will clarify the appropriate ethics review process and the institution responsible for submitting the study.
The pilot will focus primarily on lecturers as participants. Where course data are used, particular attention will be given to whether identifiable student information is accessed or recorded.
The evaluation should avoid collecting identifiable student information unless it is strictly necessary. Survey participants will also be asked not to include student names or other identifiable information in open-text responses.
The research plan will document:
- what Moodle/STACK data the plugin accesses;
- whether the plugin stores any student data;
- whether any data leave the Moodle installation;
- who can access research data;
- how data will be anonymised or pseudonymised;
- where research data will be stored;
- retention arrangements;
- participant consent and withdrawal procedures.
The experimental AI-based Model Analytics component will remain outside the initial pilot to avoid introducing additional questions concerning external processing, automated interpretation, and validation before the basic analytics workflow has been evaluated.
8. Proposed Timeline and Responsibilities
A practical sequence would be:
Phase 1 – Preparation
- finalise the plugin version for testing;
- prepare installation documentation;
- prepare short video tutorial;
- finalise research questions;
- finalise survey and pilot tasks;
- determine ethics requirements.
Phase 2 – Small pilot
- recruit approximately 5–10 lecturers;
- test installation and access;
- conduct task-based use;
- collect survey feedback;
- record technical performance.
Phase 3 – Revision
- analyse initial findings;
- correct technical issues;
- simplify or clarify confusing analytics;
- improve documentation;
- refine research instruments.
Phase 4 – Wider evaluation
- expand testing to additional universities/courses;
- collect a larger dataset;
- prepare a research paper on the evaluation.
The precise division of responsibilities between the plugin developers and research team will be agreed before the pilot begins. This should include responsibility for technical support, recruitment, ethics documentation, data analysis, and preparation of research outputs.
Expected Outcome
The immediate purpose of the pilot is to determine whether the plugin is sufficiently usable, efficient, technically reliable, and educationally meaningful to justify wider testing.
The key criterion will not simply be whether lecturers can generate analytics, but whether the plugin helps them move more efficiently from STACK data to questions such as:
Which parts of my assessment need attention?
What are students doing on this question?
What should I investigate or change as a result?
The findings from the pilot will therefore inform both continued plugin development and the design of a larger study examining the role of integrated analytics in supporting research-informed teaching with STACK.
Instructor Feedback Form
Lecturers taking part in the pilot should complete the survey referenced in § 4–6 above here: STACK Quiz Analytics Instructor Feedback Form.