Integrating AI into exams

Juggling powerful creative tools with human-led quality assurance.

Advertisement
Anaesthesia 2026 advertisement
Anaesthesia 2026 advertisement

Authors: 

On behalf of the AI Examinations Development Short Life Working Group:

  • Dr Joseph Alderman, NIHR Clinical Lecturer in Anaesthetics, University of Birmingham
  • Dr Coralie Carle, Chair, Primary FRCA
  • Dr Graeme Flett, Vice Chair, Primary FRCA
  • Dr Allan Howatson, Written Exams Lead, Primary FRCA
  • Dr James Shorthouse, Member of written exams core group
  • Dr Helen Burdett, Chair, Final FRCA
  • Dr David Leslie, SBA Lead, Final FRCA
  • Dr Cameron Weir, Examiner, Final FRCA
  • Dr Tom Billyard, SBA Lead, FFICM
  • Dr Saravana Kanakarajan, SBA Lead, FFPM

Our patients trust us as anaesthetists, intensivists and pain specialists to guide them through the most daunting and difficult days of their lives. This trust is earned through a rigorous postgraduate training programme and the dedicated study and assessment required to attain FRCA, FFICM and FFPMRCA qualifications.

Maintaining the high standards of these postgraduate examinations requires a continuous process of question generation. As we prepare for the 2027 changes to the FRCA, FCICM and FFPMRCA examinations, the demand for high-quality Single Best Answer (SBA) items will increase substantially.

To help us meet this demand for new examination questions, the RCoA, alongside the Faculty of Intensive Care Medicine and the Faculty of Pain Medicine, has established a short-life working group (SLWG) to evaluate the use of Generative Artificial Intelligence (GenAI) in the development of questions. The aim of the SLWG is to support our examiners to use these powerful creative tools responsibly by building a framework of human-led quality assurance.

Capacity and validity of examination questions

Writing high-quality examination questions is a time-consuming and cognitively demanding task. At present, each SBA question requires at least two hours of examiner time, including initial drafting and an extensive peer review and quality assurance (QA) process. To ensure our assessments remain fair and valid, new questions must be frequently cycled into examination banks, both to mitigate the impact of item leakage and to maintain curriculum coverage.

Our examiners are volunteers who fit this work around demanding clinical schedules. The intention of this proposal is not cost saving. By reducing the time required for the initial idea generation and drafting phase, we can allow our examiners to focus their expertise where it matters most – on the clinical accuracy, relevance, and educational value of the assessment.

GenAI as a creative assistant

We recommend using College-provided GenAI tools to assist examiners in the initial drafting of SBA stems and distractors. This isn’t automation – we propose a human-led workflow, where user prompts are used to instruct the GenAI tool to generate draft ideas or content based on particular curriculum points. A multi-stage QA process would ensure rigorous checking and amendment where needed. This approach offers several potential benefits over the current manual approach:

  • efficiency – by accelerating the production of workable question drafts, meaning more examiner time can be spent in editorial and QA stages
  • curriculum coverage and accuracy – helping to identify and address ‘blind spots’ and outdated or imprecise content in our question banks, ensuring that all areas of the curriculum are adequately tested
  • standardisation – generating items that are easily interpretable and adhere to formatting rules, reducing both the administrative burden of editing and comprehension for those sitting assessments.

Governance and quality assurance

Large Language Models (LLMs), which power GenAI systems, have many failings. These include:

  • ‘hallucinations’ (the fabrication of facts which may superficially seem correct)
  • sycophancy (the model agreeing with a user’s input despite evidence to the contrary)
  • the potential for algorithmic bias (where harmful tropes, outdated assumptions and other forms of discriminatory or oppressive content are embedded within algorithms, influencing their outputs).

LLM development is concentrated in a small number of large US-based multinational companies, meaning that there could be divergence in values between LLM producers and end-users. There is a risk of vendor lock-in (where end-users become dependent on a specific tool, perhaps because of lost corporate knowledge of how to function without the assistance of LLMs). Prompt injection can enable bad actors to embed malicious instructions which cause LLMs to provide inaccurate responses, or to leak information, leading to the compromise of prompts or even entire questions.

The first step in mitigating these risks is to acknowledge and understand them. The SLWG incorporates members with expertise in AI regulation and governance, and one of its actions will be to establish a rigorous QA workflow suitable for use with AI-generated questions. GenAI-derived questions won’t progress into examinations unless we’re satisfied that these issues can be effectively managed. This means that:

  1. no AI-generated item can bypass human review. Every question drafted with AI assistance is subject to the exact same peer-review process as a human-authored question. The origin of each draft question is secondary to its clinical accuracy and psychometric performance
  2. automation bias is mitigated through examiner education. Examiners using GenAI will be trained to critically appraise AI outputs, ensuring that they are primed to spot plausible-sounding but incorrect content.

Next steps

We’re not an outlier in proposing to use generative AI to assist with professional examination question setting. Several other UK colleges and faculties are conducting similar investigations or already use tools in this way. We recognise that the integration of generative AI into assessments carries significant responsibility and may prove controversial. Trust is fundamental to the examination process, and we’re committed to maintaining this through transparency. As the SLWG progresses, we’ll share our rationale and key decisions openly. We’re currently in the early stages of testing these tools, and will report on the outcomes of our technical and governance reviews in due course.

For those interested in the wider context of how we maintain examination standards, further information is available in these links:

If you have a question or comment you’d like to feed 
into the SLWG, please complete this form.