The New Default. Your hub for building smart, fast, and sustainable AI software

See now
Glossary/Artificial Intelligence

AI Risk Management

AI risk management is the discipline of identifying, measuring, mitigating, and monitoring the harms an AI system can cause.

What Is AI Risk Management?

AI risk management decides which failures of an AI system a company will accept, and who answers for them when they happen. It applies to every AI system a company builds or buys, and treats a model's errors as a managed exposure, with a named owner and controls that keep running after launch.

The risks it covers are specific to how AI behaves. A language model can state something false with full confidence, which NIST calls confabulation and most teams call hallucination. A model trained on historical decisions can reproduce the bias in them. Model drift sets in when the data a model sees in production moves away from the data it learned from, and accuracy slides without any code changing. Prompt injection, where text hidden in an input overrides the model's instructions, sits at number one in the OWASP Top 10 for LLM Applications 2025. Data leakage covers a model revealing training data, system prompts, customer records or API keys it was given access to.

Most organizations anchor the work to one of two framework families. The NIST AI Risk Management Framework (AI RMF 1.0), released on January 26, 2023, is voluntary and organizes the work into four functions: Govern, Map, Measure and Manage. Its Generative AI Profile, NIST AI 600-1, published on July 26, 2024, applies that structure to generative models and centers on 12 risks and just over 200 suggested actions. NIST states that AI RMF 1.0 is being revised under the July 2025 White House AI Action Plan; as of October 2026, 1.0 remains the current version.

On the ISO side, ISO/IEC 23894:2023 adapts general risk management guidance to AI, and ISO/IEC 42001:2023, published in December 2023, sets requirements for an AI management system that an external auditor can certify. Teams often pair them: 42001 supplies the management structure, and 23894 describes how to run the risk process inside it.

How Does AI Risk Management Protect Your Business?

  • Legal exposure under the EU AI Act. The AI Act (Regulation (EU) 2024/1689) sorts AI into four tiers: prohibited practices, high-risk systems, systems with transparency duties such as chatbots, and minimal-risk systems with no new obligations.

    Prohibitions have applied since February 2, 2025, and duties for general-purpose AI models since August 2, 2025. Article 9 requires providers of high-risk systems, such as AI that screens job candidates or scores credit, to run a documented risk management system across the system's whole lifecycle.

    The Digital Omnibus on AI (Regulation (EU) 2026/1744), in force since July 27, 2026, moved that deadline for stand-alone Annex III systems from August 2, 2026 to December 2, 2027, and for AI built into products already covered by EU safety law (Annex I) to August 2, 2028. Under Article 99, fines for prohibited practices reach €35 million or 7% of worldwide annual turnover, whichever is higher, so a company with €2 billion in turnover faces a ceiling of €140 million.

  • Failures that surface after launch and fall between teams. A model that passed every pre-release test can start misclassifying when customer behavior shifts, or reveal its system prompt to the first user who phrases a request the right way. Incidents like these sit on the line between data science and security, and each team tends to assume the other one checked. AI risk management gives every deployed model an accountable owner and monitoring that tells that owner when something changes.

How Does AI Risk Management Work?

  • Govern sets the rules and the owners. The organization defines its risk tolerance and names an accountable owner for each AI system. It also keeps an inventory of every model in use, including ones embedded in purchased software. NIST treats Govern as cross-cutting: it shapes how the other three functions run instead of happening once at the start.

  • Map establishes context for one system. The team documents what the system is for and what happens to the people it affects when it gets an answer wrong. This is also where a system gets its EU AI Act classification, since a resume screener and an internal search tool land in different tiers with very different obligations.

  • Measure turns risks into numbers and test results. Measures include hallucination rate on a fixed evaluation set, error-rate gaps between demographic groups, the share of red-team attacks that succeed and how often the model reveals data it should keep private. Red teaming means deliberately attacking the system with adversarial inputs to find failures before users or attackers do.

  • Manage decides what to do about each risk. Each identified risk is mitigated, transferred, accepted or avoided, and the decision is recorded with its owner. Mitigations range from input and output guardrails to human review of high-stakes decisions. The last resort is retiring the system.

  • Monitoring closes the loop in production. Drift detectors and incident reports feed back into Map and Measure, so the risk picture updates as the model and its data change. A model provider's version update counts as a change and triggers a fresh round of measurement.

What Tools Do Teams Use for AI Risk Management?

  • Red teaming and vulnerability scanning. PyRIT, Microsoft's open-source Python Risk Identification Tool, automates adversarial attacks against generative AI systems. garak, NVIDIA's open-source LLM vulnerability scanner, probes models for prompt injection, jailbreaks, data leakage and hallucination.

  • Evaluation and production monitoring. Evidently AI is an Apache 2.0 open-source framework that tracks data drift in classic ML models and scores LLM outputs for hallucination, PII exposure, toxicity and factual accuracy. Arize AI traces and evaluates LLM and agent behavior in production, with an open-source option, Phoenix, for self-hosting.

  • Governance and compliance platforms. Credo AI keeps a registry of an organization's AI systems and maps them to ready-made policy packs for the EU AI Act and NIST AI RMF, among other frameworks. IBM watsonx.governance combines lifecycle risk assessment, regulatory mapping, policy enforcement and model monitoring in one platform.

What Are the Key Characteristics of AI Risk Management?

  • It runs for the life of the system. It assesses risk before a model ships and monitors it in production until the model is retired. The EU AI Act writes this into Article 9, which describes the risk management system as a continuous, iterative process.

  • It counts harm to people outside the system. Traditional IT risk asks what happens to the company. AI risk management also asks what happens to the job applicant who was screened out or the patient who received a wrong summary, which is why NIST describes AI risk as socio-technical.

  • Effort scales with risk level. An internal summarization tool and a credit-scoring model both get assessed, but only the second needs bias testing across protected groups and conformity documentation. Proportionality keeps the practice affordable across a large AI portfolio.

  • Evidence is part of the output. Each decision leaves a record: test results, accepted risks, owners, and dates. That record is what an ISO/IEC 42001 auditor or an enterprise customer's security review will ask to see.

  • Responsibility spans the supply chain. Most AI products sit on a model someone else trained. The company that trained the model and the company that builds it into a product each carry part of the risk, and contracts have to say which part.

What Are the Benefits of AI Risk Management?

  • Shorter enterprise sales cycles. Large buyers now send AI-specific questionnaires during procurement. A team with a risk register and test evidence on file can answer them from existing documents instead of assembling answers per deal.

  • Fewer surprises in production. Drift alerts and scheduled re-evaluation catch accuracy loss while it is still a metric on a dashboard, before it becomes a customer complaint or a wrong decision at scale.

  • A clear path to EU AI Act compliance. Mapping each system to its tier early tells a company which products need the full high-risk treatment and which need little beyond transparency notices. The December 2027 deadline then becomes a planned project instead of a scramble.

  • Reusable controls across projects. Once a team has an evaluation harness and a guardrail layer, the next AI feature inherits them. The second and third projects cost less to bring under control than the first.

  • Better model and vendor choices. Measuring candidate models against the same risk tests gives product teams a defensible basis for choosing between a larger hosted model and a smaller self-hosted one, beyond benchmark scores and price.

What Are the Challenges of AI Risk Management?

  • Measurement needs labeled data that goes stale. Hallucination and bias rates only mean something against a curated evaluation set with known correct answers. Building one takes domain experts' time, and every product change that shifts what users ask erodes its relevance, so the set needs an ongoing maintenance budget.

  • Third-party models limit visibility. A company using a hosted foundation model sees its outputs and the vendor's documentation, little more. Pinning a specific model version keeps behavior stable for testing, but it delays the vendor's quality and security improvements until the team re-runs its evaluations on the new version.

  • Guardrails add friction for legitimate users. Input and output filters reduce prompt injection and harmful content, and they also add latency and block some valid requests. Tightening a filter lowers one risk while raising the false-positive rate that support teams hear about.

  • Shifting regulatory timelines tempt teams to wait. The Digital Omnibus bought high-risk providers 16 extra months. Deferring the work saves budget in 2026 and compresses the same testing and conformity work into a shorter window before December 2027.

  • Heavy governance slows experimentation. Running every prototype through a full review kills momentum. A lighter track for low-risk use cases restores speed, at the cost of maintaining classification rules and catching the occasional system that was tiered too low.

What Is the Difference Between AI Risk Management and Traditional Software Risk Management?

Aspect

AI risk management

Traditional software risk management

What is at risk

The behavior of a model and the people its outputs affect

Delivery of a project: schedule, budget, scope, and code quality

How failures behave

Probabilistic, so the same input can produce different outputs

Reproducible, so a defect recurs under the same conditions

When risk changes

After launch as well, as data and usage drift

Mainly when code or dependencies change

How it is tested

Statistical evaluation and adversarial red teaming

Pass/fail test cases and code review

Reference frameworks

NIST AI RMF and ISO/IEC 42001

ISO 31000 and project management methods such as PMI's PMBOK

Legal driver

AI-specific law such as the EU AI Act

Contract terms and general laws on security and data protection

FAQ About AI Risk Management

Building AI-powered AI Risk Management solutions?

Monterail's AI engineering team designs and delivers intelligent software that drives real business outcomes. Let's build together.

EXPLORE AI SERVICES