EN

What MDR IIB Means For How We Build at Delphyr

In this blog series, Delphyr Engineering, we share practical insights from building AI systems for real clinical use. Delphyr is pursuing MDR Class IIb certification for its AI platform for healthcare. In this blog, engineer Tim walks us through what that pursuit actually means for his day-to-day work. 

MDR, in a nutshell


MDR, (Medical Device Regulation) is the EU framework that classifies software used in healthcare in risk classes based on their intended use and associated risks. The classification rules consider criteria that encompass invasiveness, duration of use, and potential harm to patients or users and determine a risk class. Class IIb sits toward the higher end, reflecting the kind of clinical decision support we envision Delphyr doing. 

Compliance isn't just paperwork, it is the way of working


Say MDR and most people picture a checklist: build a feature, then tick some boxes before it ships. A stamp of approval, bolted on at the end. That's not how it actually works. MDR doesn't just check what you built. It completely changes the way of working for you to build things, and the questions you're required to ask along the way. For the engineering team at Delphyr, three ways of working are affected most: how they track where information comes from, how they handle customer feedback, and what it takes for a feature to earn the label ‘done’.  


That's not a coincidence. It connects to why Delphyr chose Class IIb instead of a lighter classification: it comes with a higher bar for evidence that the software actually works, and that evidence has to keep coming after launch, not just before it. Here's what that looks like from inside the codebase.

Customer feedback comes with a deadline


The first big shift is in how feedback from customers gets handled. If a customer flags something, Delphyr has 30 days to respond or act. Not as a nice-to-have SLA, but as a genuine obligation to have a conversation about what they raised. This is part of a bigger requirement under MDR called post-market surveillance: the obligation to keep actively monitoring how a certified product performs once it's actually in use, not just at the point of certification. 

“We co-create solutions with customers, rather than guessing.” - Tim de Boer, AI Engineer at Delphyr


In practice, that means feedback isn't something the team waits around to receive. Delphyr's products are built to make giving feedback as frictionless as possible, and the team also goes out and gathers it directly. For example by talking with clinicians about how a feature is actually working for them, not just tracking what gets reported. 

Every piece of information has to know where it came from


The second shift is in how information gets tracked. Let’s take a look at the example of using AI to check clinical guidelines. Before it can be useful to a clinician, it needs to be downloaded, stored, and converted into a format the AI model within the Delphyr platform can work with. Under MDR, every one of those steps leaves a trace: when the guideline was downloaded, how it was converted, where it lives, and which version of Delphyr's code did the converting, and when.


Why does that level of detail matter? Because guidelines get version updates, and code changes too. If it later turns out a conversion was flawed and needs fixing, the team has to be able to point back to the exact document, and the exact code, that caused the problem. So a guideline that's live in the product today isn't just "the current version" sitting there, it's the end of a chain you could, in theory, walk backward through, one document and one code change at a time.


That same principle applies far beyond guidelines. Any code that touches something MDR-regulated has to be traceable back to two things first: the risk it manages and the customer need it serves. Only once that's established does the team ask the more familiar engineering question: does it actually work?

Shipping a feature is where the real work starts, not where it ends


The third shift is about testing, and it changes what ‘finished’ even means. Under this way of working, launching a feature doesn't mark the end of the job. It's closer to a starting gun. Once something is live, the real work becomes watching closely how it performs in the real world and staying in close contact with the people using it.


Here's how that plays out in practice. It starts with the risks the team has already documented. From those risks, they define what the system actually needs to be able to do (the functional requirements). From there, they get even more specific: acceptance criteria, written in enough detail to describe exactly what the code should do, step by step. Those criteria are then turned into tests.


For a lot of software, that's a clean, predictable process. If a user logs in, the system should save that. Either it does, or it doesn't. That kind of software is deterministic: give it the same input, and it always produces the same output.


Delphyr's AI isn't that kind of software. It's built on large language models, which are probabilistic rather than deterministic: instead of following a fixed set of rules to one guaranteed answer, they generate a response based on likelihoods. Ask the same question ten times, and technically, you could get ten slightly different answers. So a single pass-or-fail test doesn't really capture whether the system works.


Instead, the team builds what's called an eval: essentially, a big, structured test made up of many realistic examples rather than one. For a requirement like "if a user asks a question in Dutch, the system should answer in Dutch," that might mean building a dataset of a thousand synthetic Dutch questions, designed to reflect situations that could realistically come up with Delphyr's customers, and checking that the system gets it right at least 95% of the time. That threshold, not a single test passing or failing, is what determines whether the behavior is considered reliable enough to ship.


Nothing here happens in isolation, either. A failing eval can trace all the way back through the same chain as everything else: to the acceptance criteria, to the customer feedback that shaped them, to the underlying risk they were meant to address.

Tying it all together


These three shifts described above, aren't really separate practices. They're one loop, and they lean on each other.


Say a clinician flags that a guideline-based answer looks off. Because of the traceability described above, the team can pull up exactly what happened: which version of the guideline was live, which version of the code processed it, and when. That reconstruction is only possible because every step left a trace.


From there, the customer feedback clock starts: Delphyr has 30 days to look into it and respond. Understanding what actually happened is what turns a vague report into something actionable. A risk that needs documenting, a requirement that needs updating, or a case that gets added to an eval dataset so the fix can be tested against realistic examples going forward, not just the one that was reported.


So traceability isn't just about explaining the past. It's what makes the feedback loop possible in the first place, and it's what feeds new cases into testing. A guideline's trace leads to a feedback conversation, which leads to a risk assessment, which leads to a test. 

The bottom line


None of this is really about regulation for its own sake. It's closer to a way of working (and not just for engineering). The same principles run through how the whole company operates: understand what people actually need, turn that into something concrete, test whether it genuinely does what was intended, and close the loop with the people using it.


MDR didn't invent that discipline. But it does make it mandatory, and that turns out to be useful in its own right: it gives everyone at Delphyr a shared, non-negotiable bar to build against, rather than a standard that has to be re-argued for every feature. For a team building AI that is designed to sit in real clinical workflows, that's less a constraint to work around than a baseline that makes the work easier: everyone already knows what good has to look like before a single line of code gets written.

Curious how this all comes together at the platform level?

Want to know why we're pursuing this approach, and what it means for how AI supports the complete care process? Request a demo to see how it works in practice.

Want to know why we're pursuing this approach, and what it means for how AI supports the complete care process? Request a demo to see how it works in practice.

Delphyr

Helping healthcare professionals reclaim their time.

Contacts

Delphyr B.V.

IJsbaanpad 2

1076 CV Amsterdam

Netherlands

Follow us

2026 Delphyr. All rights reserved.

Delphyr

Helping healthcare professionals reclaim their time.

Contacts

Delphyr B.V.

IJsbaanpad 2

1076 CV Amsterdam

Netherlands

Follow us

2026 Delphyr. All rights reserved.

Delphyr

Helping healthcare professionals reclaim their time.

Contacts

Delphyr B.V.

IJsbaanpad 2

1076 CV Amsterdam

Netherlands

Follow us

2026 Delphyr. All rights reserved.