Abstract
Large Language Models (LLMs) are being explored as support tools in healthcare and regulatory work, but it remains unclear how much specialised knowledge they hold and where they may give confident yet unreliable answers, known as confabulations or hallucinations. This talk presents a baseline study using an In Vitro Diagnostic Regulation (IVDR) question answer dataset to compare responses from four LLMs: GPT2, Pythia, Gemma2 and Llama 3. The questions cover practical regulatory topics such as legal representatives, performance studies, sampling, sponsors and post market duties. The study follows three steps. Firstly, each model was asked the same set of IVDR related questions. Secondly, the answers were manually reviewed for factual accuracy, completeness and clarity. Thirdly, responses were checked for signs of hallucination, such as unsupported claims, missing uncertainty or confident answers to questions requiring regulatory interpretation. The aim is not to replace regulatory experts or offer legal advice. Instead, the work provides a baseline assessment for deciding whether a future regulatory AI system should use prompting, retrieval or fine tuning prior to formal human review, maintaining a Human in the Loop model for Responsible AI. The presentation will share early findings, discuss gaps between general language ability and specialist regulatory knowledge, and raise questions about how AI tools should be tested before use in regulatory compliance related workflows
| Original language | English |
|---|---|
| Publication status | Published online - 14 Jul 2026 |
| Event | ORAHS 2026: Transforming healthcare systems through digital innovation and artificial intelligence - Duration: 19 Jul 2026 → 24 Jul 2026 https://www.qub.ac.uk/sites/orahs-2026/ |
Conference
| Conference | ORAHS 2026: Transforming healthcare systems through digital innovation and artificial intelligence |
|---|---|
| Period | 19/07/26 → 24/07/26 |
| Internet address |
Funding
KTP ARC Regulatory
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- AI
- Regulation
- Clinical trials
Fingerprint
Dive into the research topics of 'Benchmarking Large Language Models for In Vitro Diagnostic Regulation (IVDR) Knowledge and Hallucination Risk in Healthcare Regulation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver