Assessment
Principles and practical resources for assessment in medical education.
Good assessment is not simply a test at the end of learning; it is a planned, defensible system that gathers evidence about what learners know, can do, and are becoming. Strong assessment aligns with learning outcomes, supports feedback, and helps faculty make informed judgments about progress over time.
Programmatic Assessment
1. Weakness of the traditional approach
A single end-of-course examination is often too narrow to judge the full range of learner performance. It can overlook gradual progress and over-rely on one high-stakes result.
2. What is programmatic assessment?
Programmatic assessment combines information from multiple instruments, across time, to support defensible decisions about competence, progression and learning needs.
3. Principles of programmatic assessment
Judgments are informed by repeated evidence, multiple sources, and continuous feedback. This makes decisions more fair, transparent and clinically meaningful.
Foundation
Assessment Theories
Formative Assessment
Formative assessment is most effective when it informs learning in real time. Feedback should help teachers and learners identify what is being done well, what needs attention, and how to improve before high-stakes decisions are made.
Blueprinting
Blueprinting links assessment design to curricular outcomes. It helps ensure that the test samples the intended learning domains, matches the intended level of cognition, and distributes content fairly across the curriculum.
Objective Structured Clinical Examination (OSCE)
Video reference on OSCE design, retained from the original resource collection.
Online Assessment
Standard Setting
Measurement of Reliability
Both examiners and candidates want a test that gives a similar result on different occasions with different candidates. A reliable test produces stable, reproducible scores — a candidate assessed on two separate occasions with the same test should receive very similar marks.
Common approaches to measuring reliability include:
- Test–retest reliability
- Parallel forms
- Split-half
- Coefficient alpha
- Inter-rater reliability
Reliability is one of the most important factors in developing quality assessment questions; assessment leads should be able to show evidence of reliability so that marks can be interpreted accurately against learning objectives.
Reliability can be affected by random error — environmental, processing, classification and generalisation errors — and by bias error such as weighting, rater prejudice, the halo effect, and leniency or stringency.
AMEE Guide No. 57 discusses the psychometric theories used in assessment, including Classical Test Theory (CTT), Generalisability Theory (GT) and Item Response Theory (IRT), with the advantages and disadvantages of each.