The New Default. Your hub for building smart, fast, and sustainable AI software

See now
AI Medical Device Development: From MVP to FDA Approval

AI Medical Device Development: From MVP to FDA Approval

Piotr Zając
|   Oct 7, 2026

For AI medical devices, the evidence FDA reviews is created or lost while the MVP is being built, long before anyone files a submission. FDA defines an AI-enabled device software function as a software function that meets the device definition and uses one or more AI models to achieve its intended purpose. Where the training data came from, which records stay sealed for testing, and how the model may change after launch all start as MVP decisions. Each costs little in the first month and can force a rebuild a year later, once the round is spent. Settle them before the first training run, and the same records serve both your FDA submission and your investors' diligence.

Executive Summary

FDA has reviewed device software for decades, but the trained model inside an AI medical device raises newer questions about data and drift. FDA's AI guidance turns those questions into recommendations on data, validation, and planned changes. An evidence-grade MVP drafts that evidence from the first sprint. Early-stage teams avoid a costly rebuild by choosing data sources and a sealed test set before training. A Pre-Submission then tests those choices with FDA while changes are still cheap.

What Makes AI Medical Devices Different From Other SaMD?

AI medical devices add a second layer of review, because the model's behavior comes from training data and may shift once the device is in use. FDA reviews AI-enabled devices through the familiar 510(k), De Novo, and premarket approval (PMA) pathways, so the AI rules sit on top of existing device rules. Most AI-enabled devices reach the US market through 510(k) clearance or a De Novo grant; strictly speaking, "approval" refers to PMA. Whether a digital health product counts as a medical device at all depends on its intended use. Apps that make no clinical claim can stop there, for as long as that stays true.

For AI SaMD (software as a medical device with a model inside), the key change is the evidence behind the output. That output might be a triage priority, a flagged finding on a scan, or a risk score. FDA's draft lifecycle guidance treats the model as part of the device's mechanism of action, so reviewers examine the underlying data. The draft also flags data drift, where live deployments feed the model inputs that differ from what it learned. As of September 2026, the FDA lists over 1,600 authorized AI-enabled devices.

Why Does the MVP Stage Decide Your FDA Path?

The MVP stage decides your FDA path because the data and records a team gathers then become the evidence reviewers examine at submission. With AI, the expensive second build reaches the model and everything that feeds it.

Founders who only budget for software review may overlook the model review at submission. A team that trains a model on nearby data, then reuses it for tuning and reporting accuracy, without documenting sources or permissions, fails to meet FDA guidance. The guidance requires independent test data from different sites, which this team cannot prove. Since the evidence can't be reconstructed later, they recollect data, retrain the model, and rewrite the pipeline.

Area

MVP built for demo

MVP built as first draft of FDA evidence

Training data

Convenient data, with source and permissions unrecorded

Sources, dates, and permissions logged per batch and checked for fit with the intended population

Validation (test) data

A held-out slice of the training source, reused during tuning

A sealed test set kept from developers and drawn from sites outside the training set

Model behavior after launch

Retrained whenever it seems to help

Locked, or changed only within a written plan

What changes need a new submission

Unknown until a release is due

Mapped in advance, with significant changes outside an authorized PCCP likely requiring a new submission

Documentation

Code, a backlog, and notebooks

Dataset versions, split records, annotation rules, model version history, and a risk file

Typical founder mistake

Reporting accuracy from records that also trained the model

Collecting data before settling the intended use

Every row in the right-hand column is a recording habit during the work, and that habit is how software architecture becomes regulatory evidence.

What Does FDA Expect to See About Your Training Data?

For AI medical devices, FDA expects a documented story for every dataset. Alongside the final PCCP document, the FDA AI/ML in SaMD guidance includes a draft on lifecycle management. Its data section asks sponsors to describe each stage of that story.

  • Collection, including sites, time period, and whether any pre-existing database suits the purpose

  • The reference standard behind each label

  • Annotation, including annotator expertise and instructions

  • Storage, with version control for every dataset

  • Splits, with a record of how training, tuning, and test data stay independent

The data should reflect the intended use population, and results should be broken out by subgroup. The draft adds that relying on a single test site is generally not appropriate for assessing representativeness. EU founders should note that the draft treats data from outside the US as a potential source of bias when populations or care standards differ.

Infinant Health's discovery phase shows how early these questions surface. The company wanted a companion app for parents using its infant probiotic, with cry, stool-photo, and breastfeeding analysis planned. The proposed roadmap scoped version 1.0 as a data-collection tool to feed future ML models, with version 2.0 progressing toward SaMD certification. The app deliberately gave no medical advice before then. The discovery phase also produced a technical prototype that validated an Apple Watch-to-iPhone pipeline for on-device audio capture. When the first release exists to collect training data, the pipeline that captures it belongs in the MVP scope. 

In week one, register each dataset with its source, collection dates, and the permission for its use. Freeze a test set before any model sees it, and record who may access it. Write down the reference standard and the annotation instructions, then version datasets as carefully as code.

Locked or Adaptive Model: Which Should Your MVP Be?

For AI SaMD, a locked model is the safer MVP default, because you can test each released version on sealed data and describe it exactly. A locked model behaves identically between releases, while an adaptive model updates itself or learns from new data after launch.

FDA's final PCCP guidance covers both paths: models that engineers modify manually and models that update automatically. FDA has less precedent for models that update themselves, and it suggests setting boundaries and discussing them through a Pre-Submission.

The trade-off influences the roadmap development. A fixed model produces the clearest initial submission, but any performance improvements after launch need either a new submission or a modification plan included in the original. An adaptive design speeds up improvements by setting limits on what can be modified, how to test these changes, and what prevents a failing update. Regardless of the approach, FDA's draft guidance recommends that sponsors record differences between tested and released models and evaluate their impacts.

How Does an FDA PCCP (Predetermined Change Control Plan) Work?

An FDA PCCP, or predetermined change control plan, lets a manufacturer get specific future model changes authorized as part of a marketing submission. FDA reviews the plan with the device, so changes made under it proceed without a new submission each time.

FDA expects a PCCP to have three components:

  1. The description of modifications names each planned change and the performance it must reach.

  2. The modification protocol sets out data management, retraining, performance evaluation, and update procedures, with acceptance criteria defined in advance.

  3. The impact assessment weighs the benefits and risks of each change and of all changes together.

The plan has limits. FDA recommends a short list of specific changes that the sponsor can verify and validate. Every change must keep the device within its intended use. The plan authorized with UpDoc, an insulin-management SaMD cleared in December 2025, shows the level of detail. It lists categories of planned change, from default values to added language support, all held to the same testing framework.

Start planning when the data strategy takes shape, because the plan relies on the same data and test records. FDA encourages sponsors to get feedback on a proposed plan through the Pre-Submission program.

What Does Clinical Validation of an AI Model Involve?

Clinical validation of an AI medical device shows that the finished device performs reliably for its intended use on data the model never saw during development. FDA's draft guidance requests results for the full intended use population and for subgroups such as sex, age, race, ethnicity, disease severity, and collection site. Performance estimates should include confidence intervals. 

The draft also suggests that at least three geographically diverse US sites may suit clinical validation. Founders validating from the EU should anticipate that heavy reliance on non-US validation data may call for a higher share of US data.

Data scientists often call the tuning set a validation set. In FDA's vocabulary, validation confirms that the device meets its intended use, and the sealed set is test data.

Input variation belongs in the protocol as well. Joii's AI-powered scanner, built with Monterail's help, had to analyze photos taken under different lighting and camera angles, and the case study reports 99% image-processing accuracy. A figure like that becomes validation evidence once a protocol fixes the reference standard, the sites, and the success criteria in advance.

Pilots often change the product. A pilot puts the model in front of live users, devices, and sites, where inputs start to differ from the training data and buyers start asking what makes the AI trustworthy. For Convatec, Monterail built an Apple Watch and iOS companion app that correlates heart-rate variability with catheter usage, with risk analysis and test-protocol preparation for SaMD certification in scope. We delivered a system ready for clinical trials and helped Convatec start those trials with a different product configuration than originally planned. 

What Should You Lock Before Training Your First Model?

On a seed budget, lock eight decisions before the first training run, starting with intended use and ending with success criteria. Each becomes a record FDA may ask to see.

  1. Intended use and user. Write the claim the output supports and name who acts on it.

  2. Reference standard. Decide who or what defines ground truth and how to resolve disagreements.

  3. Data sources and permissions. List each source and the permission to use it for model development.

  4. Test set. Reserve it up front, draw it from other sites, and name who holds access.

  5. Annotation protocol. Fix annotator expertise, written instructions, and independent labeling.

  6. Model type and change plan. Choose locked or adaptive, and list the changes a PCCP might cover.

  7. Version control. Version datasets, code, and models together so you can reproduce any result.

  8. Success criteria. Set performance targets and the subgroups to report before anyone sees test results.

Each item costs a conversation or a spreadsheet, and together they give a Pre-Submission plenty to react to.

How Long Does It Take to Get an AI Medical Device to Market?

No single figure fits every device. The schedule depends on how fast you collect and annotate data, how soon validation sites sign on, which pathway applies, and how many rounds of questions the FDA sends. A 510(k) requires a predicate device, so devices with no predicate go through De Novo.

A typical path for a locked-model device going through 510(k) or De Novo looks like this. In practice, teams loop back when validation falls short, and many request a Pre-Submission more than once.

  1. MVP: record intended use, data sources, and permissions from day one.

  2. Data strategy: register datasets, fix the reference standard and annotation protocol, and seal the test set.

  3. Pre-Submission: get FDA feedback on the validation approach and any proposed PCCP before validation starts. Teams can return with follow-up Q-Submissions at later stages.

  4. Locked model: train, version, and freeze the model you will validate.

  5. Validation: test the locked model on sealed, multi-site data against criteria set in advance. A shortfall here usually sends the team back to the data or the model.

  6. Submission: file a 510(k) or a De Novo request.

  7. Post-market: monitor drift, and change the model only within an authorized PCCP or through a new submission.

FDA publishes goals for its own review clock. The MDUFA VI goals for fiscal 2027 call for a decision on 95 percent of 510(k) submissions within 90 FDA days. For De Novo requests, the goal is 90 percent within 150 FDA days, up from a 70 percent baseline under MDUFA V. These are goals rather than measured performance, and they count only FDA Days, so time spent answering FDA's questions adds to the total.

Request a Pre-Submission early. FDA's AI guidance encourages sponsors to use the Q-Submission program for feedback on development, validation, and a proposed PCCP. FDA's goal is written feedback within 70 days, or five days before a scheduled meeting, for 90 percent of Pre-Submissions.

What Changes in the EU for AI Medical Devices?

Planning for Europe adds a second regulatory system to the picture, and its outline deserves a look before the first design choices harden. The EU stacks AI medical device regulation in layers, with the EU MDR governing the device and the EU AI Act adding requirements for AI inside regulated products.

EU AI Act high-risk obligations for AI-enabled medical devices are being phased in under an evolving timeline, most recently amended by the EU's 2026 Digital Omnibus package. The practical question for a team with European plans is which of its US decisions already meet European expectations, and which need a second look.

How Monterail Approaches AI Medical Device Development Services

Most health products are built twice. Yours doesn't have to be. Monterail's HealthTech team has worked on 75+ health products, including contributions to 15+ regulated medical devices. We build to standards such as IEC 62304, with rigor matched to each component's safety class, and prepare audit-ready documentation as development proceeds. For regulatory strategy, we connect clients with MAE Group, a regulatory affairs consultancy in our Monterail Universe partner network, so that strategy informs engineering decisions early.

Remedee Labs is a machine learning example. We built native iOS and Android apps for Remedee's endorphin-stimulator wristband, using machine learning to personalize stimulation sessions from Apple Health and Google Fit data. We managed software changes so they didn't compromise the integrity of ongoing clinical trials. Remedee Labs secured CE marking and an FDA Breakthrough Device Designation.

Building an ML-powered health product of your own? Talk to our HealthTech team. We can work with you from the first prototype through validation planning to an audit-ready build.

KEY TAKEAWAYS

  • For AI medical devices, the model shapes how the device behaves, so dataset records matter as much as source code.

  • A sealed test set drawn from other sites protects the validation claim for AI SaMD, and nobody can recreate it later.

  • A locked model is the safer MVP default, and a predetermined change control plan lets planned changes proceed within written limits.

  • Eight decisions, from intended use to success criteria, belong ahead of training.

  • A Pre-Submission turns those decisions into FDA feedback, with a written response targeted within 70 days.

AI Medical Device FAQ

Author photo for Piotr Zajac
Piotr Zając
HealthTech Director
Linkedin
Piotr, Monterail’s Director of HealthTech brings over 15 years of entrepreneurial leadership and strategic innovation to the MedTech and HealthTech sectors. Piotr has demonstrated exceptional ability to build and scale healthcare solutions. Former President of EO Poland, part of the world's largest entrepreneur network. Combining his entrepreneurial background with Management 3.0 principles, Piotr specializes in helping organizations drive sustainable innovation in the rapidly evolving HealthTech landscape.