MHRA and Manchester NHS launch a health innovation sandbox: an early signal for AI-enabled devices
The Manchester Sandbox moves the evidence conversation for AI-enabled medical devices into real NHS settings. It is a regulator-enabled evidence environment, not a new authorization pathway, and key operational details have not been published.
Diana Hohage
Principal Consultant
In brief
The MHRA and Manchester University NHS Foundation Trust will test promising technologies in NHS settings with oversight and guardrails. Participation does not replace classification, conformity assessment, registration, clinical-investigation requirements or post-market surveillance.
What has been announced
The partnership brings together the MHRA and MFT. Its first major program, the Manchester Sandbox, is designed to permit the deployment and testing of promising technologies in NHS settings, with appropriate oversight and guardrails. The stated objective is to generate evidence on real-world safety, effectiveness and impact, helping the NHS to make better-informed adoption decisions and helping developers understand the evidence needed for wider NHS use.
The initial portfolio is expected to include AI-enabled medical devices. The MHRA gives as an example tools that identify people at higher risk of complications from long-term conditions, with the aim of enabling earlier support and treatment. MFT further frames the ambition in terms of earlier diagnosis, more personalized care and faster access to new technologies, while also emphasizing that learning should be shared across the wider NHS.
| Element | Confirmed position | Regulatory significance |
|---|---|---|
| Partners | MHRA and Manchester University NHS Foundation Trust | A regulator-provider partnership creates an unusually direct route to test evidence questions in a clinical-delivery setting. |
| First program | Manchester Sandbox | The sandbox is the first major program under the partnership. |
| Setting | NHS environments, involving patients and clinicians in day-to-day practice | The focus is on evidence of performance in use, not only laboratory or retrospective validation. |
| Early technologies | Expected to include AI-enabled medical devices | AIaMD developers are the most immediately relevant audience, although the program is not announced as AI-exclusive. |
| Evidence purpose | Ongoing safety, effectiveness and impact; practical evidence for NHS decision-making | The evidence agenda is explicitly lifecycle-oriented and adoption-oriented. |
| Next-phase access | MHRA will invite expressions of interest from medical-device developers | No opening date, eligibility criteria, evaluation criteria or application mechanism has yet been published. |
The program's immediate importance therefore lies less in a change to the legal rules and more in its potential to create a structured, clinically embedded setting for resolving evidence and implementation uncertainties before wider scale-up. This distinction matters: a sandbox can accelerate learning, but it cannot by itself confer regulatory status on a product.
Why the Manchester Sandbox matters now
Healthcare AI frequently performs differently after it leaves a development dataset and enters live care. Patient mix, referral pathways, local data quality, staffing, user behavior, digital infrastructure and escalation arrangements all shape whether an algorithm is clinically useful and safe. The Manchester model responds directly to this translation problem: technologies will be assessed where patients and clinicians actually encounter them, rather than only where developers can control inputs and workflows.
This approach aligns with the MHRA's wider regulatory direction. Its prior AI Airlock program was established as a regulatory sandbox for AI as a medical device (AIaMD). Phase 2 ran from April 2025 to May 2026, examined seven case studies, and generated technical and regulatory insights and recommendations. Importantly, the MHRA expressly states that the Phase 2 report does not constitute formal MHRA guidance; innovators' case studies are likewise not MHRA guidance or policy.
That caveat is not a weakness. It is the point of a sandbox. A sandbox creates evidence, tests assumptions and exposes unresolved questions. It may inform future policy, guidance, regulatory-support tools or NHS adoption practice, but its outputs should not be treated as prospective legal requirements unless and until the MHRA formally adopts them through the appropriate process.
The Manchester Sandbox appears to extend this learning logic into a regional NHS-provider setting. The release says that the collaboration offers an opportunity to test proportionate, lifecycle-based regulatory approaches fit for current and future healthcare needs. That is meaningful language, but it should be read as an opportunity for learning rather than a commitment to a particular new approval mechanism or a predetermined regulatory outcome.
What the sandbox is and what it is not
The most useful way to characterize the initiative is as a regulator-enabled, real-world evidence and implementation environment. It brings developers, clinicians, healthcare providers and the regulator closer together at an earlier stage. The goal is to generate practical evidence concerning clinical use, ongoing safety, effectiveness and impact, while identifying what would be necessary for wider NHS adoption.
It is equally important to state the boundaries. Neither of the two primary announcements says that Manchester Sandbox participation substitutes for a UKCA or recognized CE route, removes MHRA registration duties, authorizes general commercial placement on the Great Britain market, or exempts a project from applicable clinical-research and NHS-governance processes.
| Question | Correct interpretation as at 3 September 2026 |
|---|---|
| Does the announcement create a new device-authorization pathway? | No such pathway is announced. The release describes evidence generation and safe deployment and testing, not a replacement for legal market-access requirements. |
| Is the sandbox binding MHRA guidance? | No. The Manchester release is a press announcement, not guidance. The analogous AI Airlock Phase 2 report expressly disclaims formal-guidance status. |
| Is AIaMD the only eligible technology category? | No. The program concerns innovative healthcare technologies generally; AI-enabled medical devices are among the first expected technologies. |
| Are all deployment mechanics known? | No. The release does not publish project eligibility, study designs, legal routes for each deployment, endpoints, procurement terms, information-governance arrangements or timelines. |
| Can the program influence future practice? | Potentially. The MHRA and MFT say that learning could inform wider adoption and support lifecycle-based approaches, but they do not promise a regulatory change. |
For regulatory-affairs teams, the practical conclusion is straightforward: treat the Sandbox as a potential venue for evidence generation, regulator dialogue and implementation learning, not as a route around classification, conformity assessment, registration, clinical-investigation requirements, quality management or PMS.
The continuing Great Britain device-regulatory baseline
Many AI products are regulated in the United Kingdom as medical devices or in vitro diagnostic medical devices (IVDs), depending on their intended purpose and functionality. The MHRA's software-and-AI materials confirm that its responsibilities span pre-market and post-market manufacturer inquiries, technical-file review, clinical-investigation aspects, post-market surveillance and work to ensure that the regulatory framework is fit for SaMD and AIaMD.
For products placed on the Great Britain market, the MHRA's current market-regulation guidance covers certification, conformity marking and registration. All medical devices, including IVDs, custom-made devices and systems or procedure packs, must be registered with the MHRA before being placed on the Great Britain market. A UKCA pathway is available, while certain CE-marked devices may rely on time-limited recognition arrangements specified in the MHRA guidance.
Where a project constitutes a clinical investigation requiring notification, the sponsor must inform the MHRA at least 60 days before the planned start. The MHRA provides a decision flowchart and application materials for determining whether a clinical-investigation application is needed. The analysis is fact-specific; it can differ for an already marketed device used within its intended purpose, an investigational deployment, and certain in-house manufacture arrangements.
Consequently, a Sandbox candidate should map its deployment model before it sees the Sandbox as an operational opportunity. The core questions are not merely scientific. They include: What is the intended purpose? Is the product a medical device or IVD? Is the proposed use within the intended purpose? Has it been placed on the GB market? Does the protocol constitute a clinical investigation? What MHRA, research-ethics, NHS research-governance, data-protection, cybersecurity and local operational approvals are required? The Manchester announcement does not answer those questions for individual projects; each must be assessed on its own facts.
A lifecycle signal, reinforced by current PMS rules
The timing of the initiative is especially notable because Great Britain's PMS framework has become more explicit. The Medical Devices (Post-market Surveillance Requirements) (Amendment) (Great Britain) Regulations 2024 inserted a new Part 4A into the UK Medical Devices Regulations 2002. The PMS requirements cover medical devices, IVDs and active implantable devices in Great Britain and include incident-notification and preventive or corrective-action requirements after the device is first approved for the GB market.
The MHRA's detailed PMS guidance requires manufacturers to have processes for gathering and analyzing feedback and complaints, and for ensuring continued conformity with safety and performance standards. PMS data must inform risk management and, for UKCA-marked devices, technical documentation. The PMS plan must be proportionate to risk and cover comprehensive real-world data, analytical methods and links to preventive and corrective action. It also requires proactive feedback from relevant user groups, including healthcare professionals and patients where relevant, as well as consideration of usability and instructions for use.
The Manchester Sandbox is not itself a PMS program. Nevertheless, it creates a setting in which developers may be able to design a stronger bridge between pre-deployment clinical evidence, managed initial use and eventual lifecycle surveillance. The value lies in starting that bridge early rather than trying to retrofit it after a broad launch.
Lessons from the AI Airlock: useful, but not prescriptive
The AI Airlock Phase 2 report is a valuable indicator of questions that may arise in AIaMD evaluation. It should be used carefully: it informs a prudent evidence architecture, not a checklist of binding Manchester Sandbox admission requirements.
First, the report states that pre-market testing may not be sufficiently suited to demonstrate long-term, actual performance in practice. It advises manufacturers to design validation with real-world deployment in mind, identifying assumptions about user behavior, data quality and clinical context that may fail in production and specifying how those assumptions will be monitored.
Second, it treats human oversight as a lifecycle variable with regulatory significance. Human review may become less rigorous as users gain confidence in apparently reliable outputs. The report therefore suggests monitoring meaningful changes in oversight behavior, considering error severity as well as frequency, and examining how human-in-the-loop steps intersect with the device function. Review time and edit depth were discussed as possible indirect indicators of engagement quality.
Third, the report underlines the importance of intended-purpose boundaries and active guardrails, particularly for generative AI. Its TORTUS case study reported out-of-scope performance in approximately 39% of real-world notes without active guardrails and 20% with guardrails active. Those figures are case-study-specific and must not be generalized to other products; the broader lesson is that disclaimers alone are passive controls and do not reliably prevent functionality that strays beyond the intended purpose.
Finally, the report discusses predetermined change control plans (PCCPs) as a developing concept for anticipated AI modifications, while clearly noting that PCCPs are not yet formally embedded in the UK regulatory framework. It also identifies continuing uncertainty around the monitoring of changing AI systems, including what to monitor, how often, and which signals demonstrate meaningful change in performance.
| AI Airlock learning | Practical implication for a Manchester Sandbox candidate | Status of the implication |
|---|---|---|
| Real-world performance depends on user behavior, data quality and clinical context. | Pre-specify site-level assumptions, collect denominator data and define how deviations will be detected and managed. | Strong practice recommendation informed by AI Airlock; not a published Manchester criterion. |
| Human review can deteriorate over time. | Define the oversight task, escalation triggers, audit trail and indicators of meaningful review; assess error severity alongside error rate. | Strong practice recommendation informed by AI Airlock. |
| Generative systems can produce out-of-scope outputs. | Test intended-purpose boundaries; deploy active guardrails and verify them in the real workflow. | Strong practice recommendation informed by AI Airlock. |
| Change control for AI remains complex in GB. | Freeze and document model version, prompts, thresholds, integrations and workflow configuration; require documented impact assessment before any change. | Conservative regulatory practice; PCCPs are not currently an established GB pathway. |
| PMS must collect and analyze real-world feedback proportionately to risk. | Integrate structured user and patient feedback, complaints, usability data, incident assessment and corrective and preventive-action pathways. | Formal PMS concepts for applicable marketed devices; project implementation remains risk- and context-specific. |
Building a credible real-world evidence protocol
Developers considering the forthcoming expression-of-interest process should not wait for the application to build their evidence logic. A credible protocol should explain how the product's clinical promise will be evaluated in the actual care pathway and how the study will avoid confusing technical performance with clinical utility.
The proposed framework below is a professional recommendation. It synthesizes the Manchester announcement's emphasis on real-world safety, effectiveness and impact; the statutory GB PMS concepts of feedback, real-world data and corrective action; and the non-binding learning from AI Airlock. It is not a published MHRA or Manchester Sandbox template.
| Evidence domain | Questions to pre-specify | Illustrative evidence and controls |
|---|---|---|
| Intended purpose and decision claim | What clinical decision or care action does the device support? Who is the user? What is explicitly out of scope? | Stable intended-purpose statement; claims-to-endpoints matrix; boundary and prompt testing where relevant; version-controlled labeling and user instructions. |
| Clinical performance | Does the device achieve the clinically relevant performance level in the intended population and setting? | Prospective or retrospective comparator analysis; clinically meaningful reference standard; sensitivity, specificity, predictive values, calibration and clinically grounded error taxonomy where appropriate. |
| Clinical effectiveness and utility | Does use of the device improve decisions, timeliness, outcomes or resource use compared with current practice? | Predefined pathway outcomes; time-to-action; change in decision quality; avoidable investigations; relevant patient outcomes; balancing measures for workload and missed cases. |
| Human oversight | Who reviews, overrides and escalates AI output, and what happens when they disagree? | User-role design; mandatory review points; override reason capture; review-time and edit-depth analysis where meaningful; escalation protocol; targeted training and competency records. |
| Workflow and implementation | Is the device usable and safe in the local care pathway, including hand-offs and exception handling? | Usability studies; observation of workarounds; integration and latency testing; downtime procedures; training uptake; qualitative clinician and patient feedback. |
| Equity and subgroup performance | Does safety and performance vary across clinically relevant or protected subgroups, data sources or care settings? | Pre-specified subgroup plan; assessment of missingness and representation; uncertainty reporting; threshold review; fairness and access impact assessment. |
| Misuse, automation bias and out-of-scope use | How might users over-rely on, circumvent or extend the tool beyond its intended role? | Misuse scenarios; simulation; audit of overrides and out-of-scope queries; guardrail testing; review of interface cues and warnings; corrective action triggers. |
| Drift, change and configuration | How will input, output or workflow change be detected and managed? | Locked deployment baseline; model, dataset, prompt, threshold and integration configuration record; drift indicators; change-control board; revalidation rules; rollback and safe-stop process. |
| Safety, vigilance and corrective action | How are incidents, near misses, complaints and trends captured, assessed and acted upon? | Linked clinical and manufacturer reporting pathways; risk-based monitoring schedule; signal triage; CAPA process; documentation for PMS and vigilance obligations where applicable. |
| Data governance and transparency | Are data uses, access controls, provenance and disclosures clear to the relevant parties? | Data-flow map; lawful-basis and information-governance assessment; access logs; retention rules; transparency materials; cybersecurity and supplier assurance. |
A protocol should also include a clear governance charter. It should specify the sponsor, the clinical lead, the manufacturer's regulatory and quality responsibilities, the route for user feedback, the clinical-safety and information-governance functions, the decision rights for model or configuration changes, and an escalation route that can pause deployment. That governance design is central because real-world evidence becomes credible only when the people who see operational problems can convert observations into timely risk decisions.
Preparing for the expression of interest
The MHRA has said it will invite expressions of interest from medical-device developers for the next phase. At the time of writing, the announcement provides no date or published application criteria. Interested organizations should therefore monitor the official MHRA and MFT channels rather than infer an opening date or eligibility threshold.
A well-prepared applicant will be able to articulate a narrow, clinically important question rather than merely present an algorithm. The candidate should be able to explain why a Manchester NHS setting is necessary to resolve the evidence gap, what patient-safety safeguards are already mature, how the local workflow will be supported, what data will be collected, how findings will influence the product's risk management and technical documentation, and what decision would justify wider adoption.
| Preparation priority | Why it matters | Minimum readiness output |
|---|---|---|
| Device regulatory position | The legal route depends on intended purpose, device status and proposed use. | Classification and qualification rationale, regulatory-status summary and deployment-route assessment. |
| Clinical value proposition | Sandbox capacity should be directed to a meaningful unresolved clinical or operational question. | One-page claims, care-pathway and evidence-gap statement validated by a clinical sponsor. |
| Study and governance design | Real-world testing requires defined roles, endpoints, oversight and escalation. | Draft protocol, safety case, monitoring plan, governance charter and stopping rules. |
| Product configuration control | AI behavior can be affected by model, prompts, thresholds, data interfaces and workflow design. | Immutable deployment baseline, change-control log and rollback plan. |
| Lifecycle evidence plan | Evidence should support both immediate evaluation and longer-term monitoring. | Traceability from claim to endpoint to risk control to PMS and CAPA feedback loop. |
Implications for stakeholders
For manufacturers, the signal is that evidence expectations are becoming more implementation-conscious. A technically impressive validation study may be necessary but insufficient if it does not show how performance translates into a real NHS pathway, with the relevant users, data and safety controls. The productive response is to treat clinical, regulatory, quality, human-factors and health-system evidence as a single connected program.
For NHS providers and clinical teams, the Sandbox may offer a more disciplined route to evaluating technology before broad adoption. The benefit will depend on rigorous local governance: clinicians should not be positioned as a residual safety control without clear task design, adequate training, reliable escalation and feedback that actually changes the product or its use.
For regulatory and policy teams, Manchester is a useful experiment in how to make lifecycle regulation responsive to adaptive digital technologies without pre-empting the evidence. It may create a feedback loop between real-world performance, manufacturing controls and NHS adoption. However, the program's credibility will depend on transparency about project selection, governance, patient involvement, endpoints, adverse outcomes and the distinction between sandbox learning and formal regulatory policy.
Bottom line
The MHRA and MFT Manchester Sandbox is a validated and consequential early regulatory-development initiative. It is consequential because it moves the evidence conversation for healthcare innovation, including AI-enabled medical devices, into real NHS settings with the regulator and provider engaged earlier. It is early because key operational details, including selection criteria, legal deployment routes for individual projects and evidence templates, have not yet been published.
The appropriate message for innovators is neither complacency nor alarm. Participation, if available, may help generate the evidence needed for safer and more credible NHS adoption. It does not replace the underlying regulatory framework. Developers should therefore prepare an evidence package that is scientifically rigorous, clinically grounded, operationally realistic and lifecycle-ready: one that measures not just what the model predicts, but how people use it, where it fails, which patients may be disadvantaged, what changes after deployment, and how the organization will respond when the real world differs from the development environment.
This article is an information and regulatory-policy analysis, not legal advice. The regulatory route and research-governance requirements for any deployment depend on the product, intended purpose, study design, setting and facts of the individual case.
Relevant for your project?
Similar questions in your current project?
In a first call we clarify what is specifically relevant for your situation, without obligation.
Request a call →Life Science Journal
Regulatory updates, straight to your inbox.
New requirements, authority decisions and practice notes. Once a month, unsubscribe any time.
Regulations & standards considered
- MHRA and Manchester University NHS Foundation Trust, health innovation sandbox partnership, announced 2 September 2026
- MHRA AI Airlock Sandbox Phase 2 Programme Report, expressly not formal MHRA guidance
- Medical Devices (Post-market Surveillance Requirements) (Amendment) (Great Britain) Regulations 2024, new Part 4A of the UK Medical Devices Regulations 2002
- Clinical investigations for medical devices: notification to the MHRA at least 60 days before the planned start
FAQ
Frequently asked questions
Related expertise
EU AI Act →
Where sandbox learning sits in relation to binding requirements for AI-enabled devices
Software as a Medical Device →
Intended-purpose boundaries, guardrails and configuration control for deployed software
Post-Market Surveillance →
The Great Britain PMS framework that a sandbox deployment should be designed to feed
Related projects
All case studies →Sources
- MHRA, MHRA and Manchester NHS partner on health innovation sandbox, GOV.UK, 2 September 2026
- Manchester University NHS Foundation Trust, MHRA and Manchester NHS partner to bring safe healthcare innovations to patients sooner, 2 September 2026
- MHRA, AI Airlock Sandbox Phase 2 Programme Report, GOV.UK, 9 June 2026, updated 27 July 2026
- MHRA, Software and artificial intelligence (AI) as a medical device, GOV.UK, updated 3 February 2025
- MHRA, Regulating medical devices in the UK, GOV.UK
- MHRA, Clinical investigations for medical devices, GOV.UK
- MHRA, Medical devices: post-market surveillance requirements, GOV.UK, updated 5 September 2025
- MHRA, Requirements of the manufacturer's PMS system, GOV.UK, updated 5 September 2025
Related insights
All insights →Your project
Have a concrete project?
Briefly outline your situation. We'll respond with an initial assessment, usually within one business day.
Prefer direct? +49 89 4161170-0
info@theentourage.de
- Reply usually within one working day
- 4 offices: DE · CH · IT · US
- 100% life sciences



