Good assessment is not simply a test at the end of learning; it is a planned, defensible system that gathers evidence about what learners know, can do, and are becoming. Strong assessment aligns with learning outcomes, supports feedback, and helps faculty make informed judgments about progress over time.

Foundation

    Assessment Theories

      Formative Assessment

      Formative assessment is most effective when it informs learning in real time. Feedback should help teachers and learners identify what is being done well, what needs attention, and how to improve before high-stakes decisions are made.

        Blueprinting

        Blueprinting links assessment design to curricular outcomes. It helps ensure that the test samples the intended learning domains, matches the intended level of cognition, and distributes content fairly across the curriculum.

          Written Assessments

            Objective Structured Clinical Examination (OSCE)

              Video reference on OSCE design, retained from the original resource collection.

              Online Assessment

                Standard Setting

                  Measurement of Reliability

                  Both examiners and candidates want a test that gives a similar result on different occasions with different candidates. A reliable test produces stable, reproducible scores — a candidate assessed on two separate occasions with the same test should receive very similar marks.

                  Common approaches to measuring reliability include:

                  • Test–retest reliability
                  • Parallel forms
                  • Split-half
                  • Coefficient alpha
                  • Inter-rater reliability

                  Reliability is one of the most important factors in developing quality assessment questions; assessment leads should be able to show evidence of reliability so that marks can be interpreted accurately against learning objectives.

                  Reliability can be affected by random error — environmental, processing, classification and generalisation errors — and by bias error such as weighting, rater prejudice, the halo effect, and leniency or stringency.

                  AMEE Guide No. 57 discusses the psychometric theories used in assessment, including Classical Test Theory (CTT), Generalisability Theory (GT) and Item Response Theory (IRT), with the advantages and disadvantages of each.

                    Assessment of Attitude and Professionalism

                      Portfolios

                        Workplace-based Assessment

                          Logbook