Table of Contents

Model Risk Management: Framework, Process

Hao Wu
Software Engineer
|
September 3, 2026

Models turn assumptions about data and behavior into decisions about credit, pricing, staffing, inventory, and risk. When those assumptions fail or a model is used outside its design, technical error can become financial loss, operational disruption, regulatory exposure, or harm to people.

A useful program does more than validate a model before launch. It keeps an inventory, assigns accountable owners, scales review to materiality, monitors actual outcomes, and restricts or retires models that no longer perform as intended. This guide explains that operating model from initial classification through validation, monitoring, governance, and the additional controls needed for AI and machine learning.

What is model risk management?

Model risk management (MRM) is a coordinated set of policies, roles, processes, and controls for identifying, assessing, mitigating, and monitoring model risk. A model can be a statistical forecast, valuation method, credit score, fraud detector, demand predictor, optimization system, or another quantitative method that transforms inputs into estimates used in decisions.

Models are simplified representations of reality. Their usefulness depends on assumptions about data, behavior, and the environment. Model risk arises when those assumptions, the implementation, or the use of the output leads to a poor decision. A technically sound model can still create substantial risk when it is applied to a population or purpose outside its design.

The concept is closely associated with financial services. The US banking agencies' 2026 interagency guidance on model risk management defines a model more narrowly for its supervisory scope and emphasizes a risk-based approach based on inherent risk, exposure, purpose, and use. Organizations outside banking can apply the same core discipline while defining model scope around their own decisions, obligations, and risk tolerance.

Why is model risk management important?

Models can affect pricing, credit, capital allocation, staffing, inventory, fraud investigations, and customer treatment. An error can therefore propagate from a technical system into financial loss, inaccurate reporting, regulatory exposure, operational disruption, or harm to people.

MRM makes that exposure visible before a failure. A complete inventory shows where models operate. Classification identifies which models warrant deeper validation. Documentation lets reviewers understand assumptions and limitations. Monitoring detects when data or outcomes depart from the conditions under which a model was approved.

The portfolio view matters as much as individual accuracy. Several models may depend on the same vendor feed, economic assumption, customer attribute, or upstream model. A shared defect can affect them simultaneously. The 2026 interagency guidance treats these interactions and common dependencies as sources of aggregate model risk, which means an organization cannot understand its exposure by reviewing each model in isolation.

How does model risk management work?

Model risk management works as a control loop around model use. The organization identifies a candidate model, records its purpose and owner, assesses its risk, develops or acquires it under documented standards, and tests it. An appropriately independent party validates higher-risk models before approval. Production monitoring then feeds evidence back into reassessment, remediation, or retirement.

Risk determines the rigor of each activity. A model that provides a reversible internal forecast does not need the same review as one that influences a consequential customer decision or a material financial statement. Common classification factors include model complexity, uncertainty, data quality, decision impact, breadth of use, affected population, regulatory relevance, and the ability to reverse an erroneous action.

One organizing principle is effective challenge. The 2026 interagency guidance describes this as critical analysis by objective experts who have suitable expertise, enough independence to remain objective, and sufficient organizational standing to cause necessary changes. Independence alone is not enough. A validator also needs access to evidence and authority to escalate unresolved findings.

The result is a documented decision, not a binary certificate of correctness. An approval may include usage limits, overlays, monitoring thresholds, a remediation deadline, or a requirement for human review. Material changes to the model or its context send it through the loop again.

Key components of model risk management

An effective MRM program connects its core components across three control layers.

Scope, inventory, and classification. Policy defines what counts as a model, how borderline cases are assessed, and how materiality changes the requirements. A central inventory covers models under development, in use, restricted, and retired, while classification connects risk to validation, approval, and monitoring.

Development, change control, and validation. Developers document purpose, data, assumptions, methodology, testing, limitations, and intended use. Release records show what changed and why. Reviewers challenge conceptual soundness, implementation, data, performance, and use, then assign findings to owners for resolution or risk acceptance.

Monitoring, issue management, and governance. Production indicators test whether the basis for approval still holds and trigger a defined response when it does not. Named owners, approval bodies, senior oversight, and internal audit establish accountability, while portfolio reporting exposes risk across business units and shared dependencies.

These components need a common evidence trail. An inventory entry that does not connect to validation findings, versions, monitoring results, and approvals provides a catalog, but not control.

Model risk management framework

A model risk management framework translates policy into repeatable decisions. It should specify the unit being governed, the risk taxonomy, required evidence, decision rights, escalation paths, and lifecycle triggers. The framework should also distinguish model ownership, independent review, and assurance over the program itself.

The following structure works across many organizations:

Framework Layer Question It Answers Typical Evidence
Policy What is in scope, and what standards apply? Model definition, risk appetite, tier requirements
Portfolio Where is model risk concentrated? Inventory, classifications, shared-dependency analysis
Model Is this model fit for its intended use? Development record, validation report, limitations
Decision Who may approve use and accept residual risk? Approval, conditions, exceptions, remediation plan
Runtime Is the model still operating within approved bounds? Monitoring results, incidents, overrides, change records
Assurance Does the MRM program work as designed? Internal audit results, issue aging, coverage reporting

The framework should be proportionate rather than uniform. The 2026 interagency guidance explicitly notes that practices appropriate for one banking organization or model may be inappropriate for another. Proportionality does not mean leaving lower-risk models unrecorded. It means using a lighter set of controls while retaining ownership, visibility, and triggers for reassessment.

Types of model risk

Model risk has several sources, and a single failure can involve more than one.

Conceptual and data risk. The theory, assumptions, variables, or method may not represent the problem well enough. Training, calibration, or scoring data may be incomplete, stale, biased, incorrectly transformed, or inconsistent with the intended population. A weak proxy or an upstream data change can affect every result built on it.

Implementation, performance, and use risk. Code, configuration, feature logic, or integration can differ from the approved design. Performance can deteriorate or conceal failures for important segments. People can also apply an accurate model outside its intended scope, misunderstand uncertainty, or automate a decision that required judgment.

Third-party, concentration, and dependency risk. A vendor model or data service may change with limited notice, expose little technical detail, or restrict testing. Several models may also share vendors, data, assumptions, code, or upstream outputs, allowing one failure to create correlated errors across the portfolio.

These categories interact. Effective controls connect design and validation evidence to data lineage, production use, and portfolio dependencies so that one symptom is not mistaken for the complete failure.

Model risk management lifecycle

The MRM lifecycle begins before development and continues until dependencies and required records have been resolved after retirement.

Steps 1 through 3 establish scope and build the model.

  1. Intake and purpose definition. State the supported decision, intended users, affected parties, expected benefit, prohibited uses, and owner. Decide whether the proposed system meets the organization's model definition.
  1. Risk classification. Evaluate inherent uncertainty and materiality. Assign a tier that determines documentation, testing, validation independence, approval authority, and monitoring depth.
  1. Development or acquisition. Select data and methods, document assumptions, test alternatives, implement the model, and record limitations. For a vendor product, gather enough evidence to assess fitness for the organization's use.

Steps 4 and 5 establish whether the model may operate and under which controls.

  1. Validation and approval. Challenge the design, implementation, performance, and use. Resolve material findings or place explicit restrictions around them. Record the residual-risk decision and approver.
  1. Deployment and monitoring. Confirm that the approved version and controls reached production. Monitor data, behavior, outcomes, overrides, and incidents against predefined thresholds.

Steps 6 and 7 respond to change and close the model's obligations.

  1. Change and reassessment. Reclassify and revalidate when the model, data, purpose, population, vendor, integration, or operating environment changes materially.
  1. Retirement. Stop use, revoke access, update downstream dependencies, preserve required evidence, and confirm that another process has not continued to consume the retired output.

The lifecycle is a feedback loop. Monitoring and material changes return evidence to classification, validation, and governance, while retirement closes both the model and its downstream dependencies.

Figure: Model risk management is a feedback loop: production evidence returns to classification and validation, while shared governance records preserve accountability through retirement.

Model inventory and classification

A model inventory is the control plane for the program. It should include models being developed, used internally, purchased from vendors, embedded in applications, restricted, and retired. Stable identifiers prevent names and ownership changes from breaking the evidence trail.

Useful fields include purpose, business and technical owners, status, risk tier, methodology, version, implementation location, input data, upstream and downstream dependencies, affected processes, validation status, limitations, monitoring plan, vendor, approvals, and next review trigger. The record should also identify non-model tools that were assessed and excluded, so the scope decision can be revisited when they change.

Classification should combine inherent risk with materiality. Inherent risk reflects characteristics such as complexity, assumptions, data constraints, and uncertainty. Materiality reflects the purpose and exposure of the model, including the consequence and breadth of decisions influenced by its output. Controls then scale to the combined exposure.

Avoid reducing classification to a score whose meaning is unclear. A small number of tiers with explicit control consequences is easier to operate. For example, a higher tier might require validation before first use, approval by a designated committee, tighter monitoring, and stronger limits on unresolved findings. The rationale matters more than a mathematically precise label.

Model validation and testing

Model testing is a development activity that evaluates whether a model performs as intended and may be conducted by developers with model users. Validation evaluates performance, reliability, and limitations. The activities overlap, while governance assigns responsibilities and addresses conflicts of interest.

Conceptual soundness. Review the purpose, theoretical basis, variable selection, assumptions, methodology, data relevance, and limitations. The central question is whether the design makes sense for the intended decision.

Process and implementation verification. Confirm that code, calculations, transformations, interfaces, access controls, and production configuration implement the approved design. Reproduce key results where practical and test failure handling.

Outcomes analysis. Compare predictions or estimates with observed outcomes using measures aligned with the model objective. The 2026 interagency guidance identifies back-testing and outlier analysis as possible forms, while leaving the approach dependent on methodology, objectives, and data availability.

Benchmarking and sensitivity analysis. Compare the model with credible alternatives or simple baselines. Vary important assumptions and inputs to find brittle behavior, discontinuities, and conditions under which conclusions change.

Use and control review. Determine whether users understand limitations, whether overrides are recorded, and whether surrounding controls prevent use outside approved bounds.

Validation depth should reflect risk, but every report should state the scope, limitations, findings, required remediation, and whether the model is fit for its proposed use. Conditional approval should name the condition and its owner.

Model monitoring and performance management

Monitoring tests whether the basis for approval still holds in production. It needs metrics for the model, its data, and the business process around it.

Input monitoring can track missingness, range violations, category shifts, population drift, source freshness, and schema changes. Output monitoring can track score distributions, confidence, stability, and unusual concentrations. Outcome monitoring measures relevant accuracy, calibration, financial results, error costs, subgroup behavior, or other real-world effects once labels become available.

Operational evidence matters too. Track overrides, fallbacks, complaints, incidents, latency failures, policy exceptions, and use outside the approved population. A model can retain predictive accuracy while creating risk because users have changed how they act on its output.

Each metric needs an owner, cadence, threshold, and response. A warning may prompt analysis; a severe breach may restrict use, increase human review, revert to a prior version, or suspend the model. Thresholds set only after a breach invite inconsistent decisions.

Performance management closes the loop through issue resolution. Teams should record the cause, affected versions and decisions, interim controls, remediation owner, due date, and evidence of closure. Monitoring should also confirm that remediation worked rather than treating a deployed code change as proof.

Model governance and controls

Governance determines who owns model risk and who can act on it. The model owner is accountable for appropriate use and performance. Developers produce the model and its evidence. Validators challenge it. Business users provide context and observe outcomes. A model risk function maintains policy and portfolio reporting. Internal audit assesses whether the program operates as designed without duplicating routine validation.

Clear separation helps manage conflicts of interest, especially for material models. Organizational charts alone do not create effective challenge. Validators need appropriate expertise, access to developers and evidence, and an escalation route that cannot be neutralized by delivery pressure.

Controls should cover model changes, access, approvals, exceptions, vendor updates, issue aging, and retirement. Senior reporting should show more than a count of models. Useful views include high-risk models with overdue findings, unvalidated use, threshold breaches, shared dependencies, material vendor concentrations, and accepted risks approaching expiry.

These questions are relationship-heavy. A reviewer may need to move from a failed data source to every model version, business process, validation, exception, and owner that depends on it. PuppyGraph lets teams define that graph over records already held in SQL databases, data warehouses, and data lakes or lakehouses, including direct reads of Iceberg and Delta Lake tables. On its default direct-query path, the underlying records stay in their governed sources. Teams can use openCypher or Gremlin to trace shared dependencies and find aggregate exposure without building a separate graph-specific data pipeline. Policies still define authoritative records, access, retention, approval rights, and accountability for incomplete data.

Model risk management for AI and machine learning

Machine learning models fit the MRM lifecycle, but they expand the evidence needed. Data represents behavior rather than merely supplying inputs. Training can be stochastic. Performance may vary across populations, and explanations may be less direct. A complete review therefore covers data provenance and representativeness, feature and label construction, leakage, uncertainty, subgroup behavior, robustness, reproducibility, and the full application around the model.

Context remains decisive. A highly complex classifier used for a low-stakes suggestion may present less risk than a simple regression used in an important eligibility decision. NIST AI RMF 1.0 provides a voluntary, cross-sector framework organized around Govern, Map, Measure, and Manage. NIST describes those functions as continuous rather than an ordered release checklist. They can extend an existing MRM program to impacts on individuals and society, trustworthiness characteristics, human oversight, and risks from third-party AI components.

Generative AI requires system-level review. The model may be combined with prompts, retrieval sources, filters, tools, memory, and human review. Evaluating the base model alone does not establish that the application is safe or fit for purpose. The NIST Generative AI Profile identifies risks that are unique to or exacerbated by generative AI and maps suggested actions to the AI RMF.

Evaluation should use realistic tasks, user roles, languages, retrieval conditions, and foreseeable misuse. Relevant measures may include factual accuracy, groundedness, policy violations, harmful bias, privacy leakage, security behavior, and whether human reviewers catch consequential errors. Because outputs can vary, preserve prompts, component versions, test data, sampling settings, and evaluation methods well enough to reproduce the decision.

Agentic systems add actions and tool permissions to model risk. Govern the full loop from input and planning through tool calls, approvals, side effects, and logs. Enforce least privilege outside the model, require approval before consequential or irreversible actions, and retain a reliable way to pause or revoke access. These controls complement model validation because a valid model output can still cause harm when an application grants it excessive authority.

The 2026 interagency guidance on model risk management explicitly excludes generative and agentic AI from its scope while stating that organizational risk management and governance should guide appropriate controls for tools outside the document. That boundary is a reason to connect, not conflate, traditional MRM and broader AI governance.

Conclusion

Model risk management turns uncertainty about models into visible ownership, evidence, limits, and response paths. Its foundation is a complete inventory, risk-based classification, disciplined development, effective challenge, validation, production monitoring, and governance that considers both individual models and shared dependencies.

The framework should follow the consequence and context of use. Keep the lifecycle active as data, models, vendors, users, and business conditions change. For AI systems, widen the unit of review to include the application, affected people, human oversight, and any tools the model can operate.

Try the forever-free PuppyGraph Developer Edition and book a demo with the team to see how openCypher and Gremlin queries trace model, data, control, and owner dependencies across warehouse and lakehouse tables, with no graph-specific ETL, so teams can investigate aggregate model risk against current governance records.

Hao Wu
Software Engineer

Hao Wu is a Software Engineer with a strong foundation in computer science and algorithms. He earned his Bachelor’s degree in Computer Science from Fudan University and a Master’s degree from George Washington University, where he focused on graph databases.

Get started with PuppyGraph!

PuppyGraph empowers you to seamlessly query one or multiple data stores as a unified graph model.

Dev Edition

Free Download

Enterprise Edition

Developer

$0
/month
  • Forever free
  • Single node
  • Designed for proving your ideas
  • Available via Docker install

Enterprise

$
Based on the Memory and CPU of the server that runs PuppyGraph.
  • 30 day free trial with full features
  • Everything in Developer + Enterprise features
  • Designed for production
  • Available via AWS AMI & Docker install
* No payment required

Developer Edition

  • Forever free
  • Single noded
  • Designed for proving your ideas
  • Available via Docker install

Enterprise Edition

  • 30-day free trial with full features
  • Everything in developer edition & enterprise features
  • Designed for production
  • Available via AWS AMI & Docker install
* No payment required