The CSBS AI Supervisory Framework: What It Means for Mortgage Servicers and Bank Servicers
State examiners now have a common playbook for finding AI inside the institutions they supervise. For servicers, the framework matters more than its modest packaging suggests, because servicing is the most embedded-AI-dense operation mos...

The framework is a discovery tool, and servicing is where it will find the most
State examiners now have a common playbook for finding artificial intelligence inside the institutions they supervise and deciding how closely to look at it. The Conference of State Bank Supervisors approved its Artificial Intelligence Supervisory Framework on August 13, 2026 and released it publicly in September. It creates no new legal obligation, and every state agency decides for itself how much of it to use. None of that makes it safe to ignore.
For servicers the framework matters more than its modest packaging suggests, because servicing is the most embedded-AI-dense operation most state-regulated institutions run. The AI in a servicing shop mostly arrived inside things the servicer bought rather than things it built. It sits in the contact center platform and in the document intake tool. It runs inside the decision engine behind the modification waterfall, and increasingly it ships inside the system of record itself. An institution that has never written the words “artificial intelligence” into a policy can still have a dozen AI use cases touching borrowers every day.
That is why the framework opens with discovery rather than with standards. Its first scoping question asks whether the institution uses AI anywhere in its business, and it instructs examiners to test a negative or unclear answer against the institution’s vendor and software inventories and against recent platform changes before accepting it. The servicers who have the hardest time under this framework will be the ones who answer that first question with “we don’t use AI.”
What CSBS released and how it reaches a servicer
CSBS released the framework publicly on September 16, 2026, after its State Supervisory Processes Committee and its Nondepository Supervisory Committee both approved it in August. That second approval matters for servicers. The nondepository side of CSBS signed off, so the package was built with licensed nonbank institutions in view and not only state-chartered banks.
The package has five parts. The Core Examiner Guide is the spine. It opens with eight scoping questions and a document request list, works through the three review sections shown in the map below, and closes by routing examiners to the supervisory resources that already exist. The Examiner Work Program runs 28 pages and gives examiners the purpose and source support behind each procedure, along with where to focus. The Nonbank AI Supplements overlay AI-specific questions onto three review areas nonbank examiners already work in. The Risk Tiering Worksheet is the one piece written for institutions to complete themselves, and CSBS describes it as an optional industry tool. A one-page list of sources rounds out the set.
Everything in it is discretionary. Each state agency decides how much of the framework enters its own program, and CSBS does not publish a count of which agencies have adopted it. That does not mean it will be rare. State agencies supervise about 79 percent of FDIC-insured banks and savings institutions, 3,355 of 4,233 as of June 30, so most bank servicers will meet these questions at a state exam sooner or later. A licensed nonbank servicer is licensed in every state where it services loans, which means it will meet the framework unevenly. One state will open its next exam with the eight scoping questions while another will not mention AI for two more cycles.
The right response to that unevenness is to build the evidence file once, to the framework, and hand the same package to every regulator who asks.
Where an examiner will find AI in a servicing operation
The scoping section tells examiners not to take “no” at face value, and the Work Program explains why: AI use can be missed when it is embedded in software or treated as a feature rather than as a distinct system. That describes servicing exactly. The AI is mostly inside things the servicer bought, and the examiner knows it.
The contact center is the densest area. Voice bots and conversational IVR handle first contact while agent-assist tools surface guidance to representatives during live calls. Speech analytics score those same calls for compliance after the fact, and portal chat answers borrower questions about escrow and payoff without a person in the loop. Each of those is customer-facing, so scoping question IS-3 is yes. Nearly all of them are vendor-provided, so IS-4 is yes. All of them run on nonpublic personal information, so IS-8 is yes. One chatbot lights up every section of the framework at once.
The default side holds a second cluster. Intelligent document processing reads hardship packages and sorts the pages. A decision engine runs the modification waterfall. Propensity or roll-rate scoring decides which delinquent borrowers get called first and how often. The back office holds a third cluster, where complaint intake gets auto-categorized and response letters get drafted, payment anomaly screening flags suspicious activity, and property preservation photos get machine-reviewed before a claim is filed.
On top of all that sits staff use of general-purpose generative tools for drafting and summarizing. Underneath it sits the system of record itself, whose vendor now ships AI features inside releases that the servicer never separately approved and may not have noticed.
The Risk Tiering Worksheet includes a system-type checklist that works well as a forcing device for this inventory. Two of its categories deserve attention. A “rules-based system with AI components” captures the hybrid decision engines that servicers have historically treated as configuration rather than models. An “agentic” system is one that takes action or completes work with limited human direction, and the framework treats it as its own category from the first page.
Generative and agentic AI get their own workstream
The federal banking agencies left a gap that the states have now stepped into. When the Federal Reserve and the OCC and the FDIC replaced their model risk guidance in April 2026, a footnote placed generative and agentic AI outside the scope of the revised guidance on the grounds that those models are novel and rapidly evolving. The Core Examiner Guide says the opposite in effect. It gives generative AI its own section with eight procedures and tells examiners to use the framework precisely where existing model risk resources may not fully address generative or agentic capabilities.
For a servicer the eight procedures map onto questions that are easy to ask and uncomfortable to answer. Which generative tools are in use, and is each one public or private, internally deployed or vendor-provided? What does each one do, and does it touch borrowers or inform decisions? What controls sit over the data staff enter into them? Who reviews the outputs before they go anywhere? The data question is the sharpest one in servicing, because a representative who pastes a borrower’s hardship letter into a public tool to summarize it has moved nonpublic personal information outside the institution’s control. GE-4 asks about that directly, and the document request list asks for the contract terms that say whether a vendor can retain or train on that data.
The framework also expects the servicer to have named the risks it is managing. The guide calls out inaccuracy and hallucination alongside prompt injection and data exposure, and it asks for a process that monitors and restricts generative use as capabilities change. A policy written in 2024 that approved one tool for one purpose will not satisfy GE-7 if the tool has since added agent features and the policy has not moved.
GE-8 is the procedure for AI that acts. Where a system can take actions with limited human direction, the examiner will ask how the institution defines what the system is permitted to do and where a human checkpoint sits in the path. The examiner will also ask whether every action is logged and reversible and whether someone can restrict or halt the system. Any pilot that lets an agent touch the system of record, whether it posts a payment or moves a file through loss mitigation, needs those four answers written down before the pilot grows. The industry is not ready for this part. In a Wolters Kluwer survey earlier this year, 72 percent of 230 bankers named kill-switch protocols or regulatory reporting of AI failures as the risk area where they were least prepared. Servicers have no reason to assume they are ahead of that curve.
The tiering worksheet has a servicing problem built in
The worksheet is the part of the framework a servicer will actually fill out, and it contains a trap that is specific to servicing. It rates each use case on four factors. Two concern consumer impact and human oversight, and two concern the harm an error or outage could cause and the sensitivity of the data involved. The highest-rated factor generally sets the preliminary tier. A final tier may differ, but only with documented rationale covering the risk drivers and the compensating controls, and any downward adjustment has to be reviewed and approved through the institution’s governance process.
The data factor rates any processing of nonpublic personal information as high risk. Essentially everything a servicer does runs on loan-level borrower data. A literal reading therefore puts the large majority of servicing use cases at preliminary Tier 3 before anyone has looked at consumer impact or human oversight at all. The worksheet also flags data elements that could serve as proxies for protected characteristics, and servicing files are full of them, from property geography and language preference to military status and the borrower’s own hardship narrative.
Two consequences follow. The governance approval step for down-tiering becomes the most exercised control in the whole program, and examiners will read those rationales closely because the worksheet tells them to. A servicer should design that approval path before tiering a single use case rather than discovering it as a pile of exceptions. And the Tier 3 control set becomes the default expectation unless the servicer can justify something lighter. That set starts with independent validation by a qualified party, which few nonbank servicers have in place for anything. It adds ongoing monitoring of fairness and consumer outcomes. It expects consumer disclosures where decision-making is automated and an incident response procedure written for AI failures, and it expects the board or an equivalent oversight body to hear about the use case and its issues.
The human oversight factor contains a detail that anyone with a quality control background should read twice. Exception-based review rates moderate and fully automated rates high, and the worksheet defines fully automated to include any situation where human review occurs only after the action has already affected the customer. Post-hoc QC sampling is therefore not human oversight for tiering purposes. It remains a valuable control, but it does not move a use case out of Tier 3 on this factor. The consumer impact factor works the same way for an automated eligibility decision on a modification or a fee waiver, which rates high the moment no person reviews the recommendation before it takes effect.
The worksheet is optional, and CSBS says so. A servicer can tier its use cases with another methodology. Whatever methodology it uses, the examiner will compare the result against these four factors, so the practical choice is to adopt the worksheet and spend the effort on the rationales rather than on an alternative scoring model.
Vendor oversight: a managed dependency, not a black box
The third-party supplement tells examiners to focus on whether the institution treats vendor-provided AI as a managed dependency rather than a black box, and it warns that an institution relying primarily on vendor representations may need additional follow-up. For a servicer that language points straight at the system of record and the contact center stack, the two relationships where the servicer has the least visibility and the most exposure.
The supplement’s contract considerations read like a checklist for amending servicing platform agreements. Does the agreement say whether borrower data can be retained or used to train or tune the vendor’s models? How does the servicer learn of a material model update, and must the vendor disclose incidents? Can the servicer get access to testing or audit information, and can it pause or revert a feature that misbehaves? What happens on termination, and is there a fallback? Most servicing platform contracts predate every one of those features and have been amended in practice by release notes. A servicer that cannot produce the contract language will be asked how it knows what the vendor is doing with its borrowers’ data.
The supplement also asks about ongoing monitoring rather than upfront diligence alone, and about output oversight where the vendor supplies recommendations. A SOC report tells an examiner about the vendor’s control environment. It says nothing about whether the servicer has ever benchmarked the vendor’s loss mitigation recommendations against its own outcomes, and that is the question the supplement is asking.
Servicing structures add a layer the supplement’s sources anticipate. A master servicer or MSR owner that uses a subservicer owns the subservicer’s AI as its own third-party risk, and the subservicer’s vendors become fourth parties. The CRI framework language cited in the supplement names fourth-party risk explicitly, so a servicer cannot answer the vendor questions by pointing at its subservicer. Concentration gets a mention too, and for a servicer the single platform that runs the portfolio is the obvious concentration point. The supplement asks what the institution would do if a critical AI vendor changed its terms or degraded in performance or became unavailable, and very few servicers have a written answer for the platform that holds every loan.
Model risk review for institutions that never had a model risk program
The model risk supplement contains the sentence that will matter most to nonbank servicers. It asks whether model risk concepts may be relevant even if the tool is not formally categorized as a model. Most nonbank servicers never built a model risk program because nothing required one, yet they run model-like things every day. A collections score sets contact cadence. A decision engine runs the modification waterfall. A screen flags suspicious payments before they post. Each of those influences a borrower outcome, and the supplement says the absence of the word “model” on an org chart does not take it out of scope.
The supplement’s drift concern is unusually real in servicing. Portfolio composition changes with every transfer, so a scoring tool tuned on one book is applied to another with no notice. Disaster forbearance waves shift delinquency behavior overnight. Contact-strategy tools trained on which borrowers answered the phone create exactly the feedback loop the supplement names, because the tool learns to call the people who already answer. The supplement asks whether ongoing monitoring addresses shifting inputs and degradation over time, and most servicers monitor the vendor’s uptime rather than the tool’s behavior.
What the framework asks for is validation tailored to actual use and risk, scaled to materiality, rather than a bank-style program imported wholesale. That is the right pitch for a mid-size nonbank. A proportionate program documents intended use and known limitations for each model-like tool. It tests the tool against the servicer’s own outcomes before and after deployment, and it keeps a human challenge step wherever the output drives a consequential action. The supplement is explicit that an institution relying on vendor representations about performance, without its own assessment, should expect follow-up.
Bank servicers have a different reference point. The revised interagency guidance in SR 26-2 replaced SR 11-7 in April 2026 with a risk-based approach tied to model materiality, and it keeps the long-standing principle that using a vendor does not shift responsibility for model risk off the institution. Two features of that guidance push bank servicers toward the state framework anyway. The federal guidance says it is expected to be most relevant to institutions above 30 billion dollars in assets, which excludes almost every state-chartered bank, and it leaves generative and agentic AI outside its scope. The CSBS framework draws no size line and covers both.
Consumer protection: specific reasons, fair servicing and the words borrowers see
The consumer protection supplement is where the framework gets closest to the daily work of a servicer. Its first consideration asks whether AI is used in customer-facing activities or in activities that support or influence decisions affecting consumers. In servicing the answer is yes in both directions at once, because the same loss mitigation workflow that produces a decision also produces the letter that explains it.
The “specific reasons” consideration lands directly on existing law. Regulation X requires a servicer that denies a borrower for a loan modification option to state the specific reasons for the denial, and the framework’s own sources list cites the adverse action provisions of Regulation B. Where AI informs that denial, the supplement asks whether the institution can identify and communicate specific and accurate reasons for the outcome, and its examiner focus explicitly treats vague or inaccurate reasoning used in place of actual factors as a trigger for follow-up. A score is not a reason. A model output label is not a reason. An AI-informed denial has to resolve to the borrower’s actual facts, and the servicer has to be able to show how it did. Whatever direction federal guidance on algorithmic adverse action takes, the states are carrying that expectation forward through this document.
The “decision influence” consideration targets tools described as support only, and it asks whether the institution understands how AI outputs affect consumer treatment even when a person signs off. An examiner will ask how often staff override the recommendation. If the answer is almost never, the tool is making the decision whatever the policy says, and the tiering changes accordingly.
Outcome monitoring creates a tension the framework does not resolve. The supplement asks whether the institution monitors for differences in outcomes across customer groups, and the Tier 3 controls expect ongoing monitoring of fairness and consumer outcomes. Servicers do not hold demographic data, so fair servicing analysis of modification approval rates or foreclosure referral timing depends on proxy methods, and the worksheet treats proxy data as sensitive. That is not a reason to skip the monitoring. It is a reason to design it deliberately, with the method documented and the data handled under the same controls the worksheet expects for sensitive inputs.
The consumer communications consideration covers AI-drafted letters and chatbot answers, and here the servicing risk is plain. A chatbot that invents a forbearance term or misstates a payoff figure has created a UDAAP problem before anyone calls it an AI problem. The supplement asks whether generated or AI-assisted content is reviewed for accuracy and for consistency with law and policy, and the document request list asks for samples of recent AI-generated consumer-facing outputs, including chatbot transcripts. A servicer that cannot retrieve those samples with evidence of review has an answer the examiner will not like.
Finally, the Tier 2 controls expect the complaint process to capture AI-related issues. That means the complaint taxonomy needs an AI tag now, so that root-cause work can tell an examiner whether a pattern of complaints traces to a tool rather than to a person.
Bank servicers and nonbank servicers face different gaps
For a state-chartered bank with a servicing operation, the architecture already exists. A third-party program built on the interagency guidance sits alongside a model risk program that now maps to SR 26-2. The compliance management system is mature, and the board already receives risk reporting. The gap is coverage rather than structure. Servicing tends to be treated as back office, so enterprise AI governance committees focus on origination and fraud while the servicing platform’s embedded features never make the inventory. Contact center AI gets bought by operations as a telephony upgrade. Generative tools get used inside servicing without anyone mapping that use to the enterprise policy. The framework will reach these banks through the state exam they already have on the calendar, often conducted jointly with the FDIC or the Federal Reserve, as a set of questions layered onto reviews that are already scheduled. National banks and federal thrifts sit outside state jurisdiction, but since the framework is sourced almost entirely from federal guidance, the expectations converge regardless of charter.
For a licensed nonbank servicer the gap is foundational. Vendor management is typically oriented toward SOC reports and financial condition rather than model behavior. There is usually no model risk function at all. There is often no written AI policy, and generative tool adoption has run well ahead of governance. The nonbank supplements were written for exactly this institution, which is why they read as overlays on vendor review and model risk and consumer protection rather than as a standalone program: the examiner expects to find those three review areas already in place and will add the AI questions to them.
Nonbank servicers do have one structural advantage. They already live under GSE and Ginnie Mae vendor oversight expectations and under investor and agency reviews that ask many of the same questions in different words. An AI inventory built to this framework does double duty as support for those reviews. New York licensees should also note that NYDFS issued guidance on September 10, 2026 on how cybersecurity risk assessments should account for AI and other emerging technologies, so the same inventory will be asked for twice, in two vocabularies, by two parts of the same regulator.
What a servicer should do now
Start with the eight scoping questions as a self-assessment, answered against the vendor inventory rather than from memory, because that is exactly how the examiner will test the answers. The output of that exercise is the AI inventory the framework asks about in its first inventory procedure. It should name a business owner for each use case and record what the tool does and who relies on its output, since those are the fields the Core Guide tells examiners to look for. It should also distinguish internal use from customer-facing use and decision support from decision-making, because that distinction drives everything that follows.
Then complete one worksheet per use case, having designed the down-tiering approval path first so that the Tier 3 inflation problem gets handled as policy rather than as a pile of exceptions. Keep the rationales short and factual. An examiner reading a down-tiering rationale wants to see the compensating control named and evidence that it operates, not an argument about why the data factor is unfair to servicers.
Governance comes next, and it does not need to be elaborate to satisfy the framework. Someone has to own AI oversight by name. A written policy has to cover generative tools and name the ones that are approved. Management has to see periodic reporting on AI use and issues, and the board or an equivalent body has to see it for the use cases that warrant it. The guide is explicit that governance should be proportionate to the institution’s size and risk profile, so a mid-size servicer with a one-page policy and a quarterly report that actually gets read is in better shape than a large one with a committee charter nobody follows.
Use the Document Request List as the readiness checklist. Two items will take longer than the rest. The first is contract language addressing whether vendors can use or retain or train on borrower data, which usually means going back to the vendor. The second is retained samples of AI-generated consumer-facing outputs with evidence that a person reviewed them, which usually means a retention process that does not exist yet. Start both early.
Where any automation acts on the system of record with limited human direction, write down the GE-8 answers now, before the pilot grows. Ask every technology vendor for an AI disclosure that explains where models sit in the product and what data they touch, along with how changes get communicated. Vendors will get this request from every client within the year, and the ones who have an answer ready will stand out.
And read the Examiner Work Program. It is the longest document in the package and the one most servicers will skip, and it is the one that shows what the examiner is actually looking for behind each procedure.
Closing thoughts, and how BlackWolf can help
Strip away the structure and the framework asks a servicer to prove four things. It has to know where AI operates in its business, and a named person has to be accountable for each use. The risks have to be assessed in a way the servicer can defend, and the controls have to match the risk. None of that is new thinking for a well-run servicing operation. What is new is that state examiners now have a shared script for asking, and the discretion to ask it of an institution at any size.
Two cautions are worth carrying into the work. Nothing in the package defines artificial intelligence; the sources document points to the Treasury lexicon for terminology, so the line between rules-based automation and AI will be drawn by individual examiners. A servicer should draw that line first in its own inventory and document the reasoning. And because every piece is discretionary, the variance across states will be wide for a while. The servicers who treat the framework as a floor will have an easier time than those who wait to see which states enforce it.
BlackWolf Advisory Group works with servicers on exactly the operational and compliance questions this framework raises, from vendor oversight and quality control to the governance that ties them together. Our AI Readiness Review walks a servicer through where AI is operating across the business, how each use is governed and controlled today, and where the gaps sit against what state examiners now expect to see. If your next exam cycle is on the horizon, or you simply want to know what an examiner would find before one does, reach out to Mirza Hodzic at mhodzic@blackwolfadvisory.com.
Sources
The framework and all of its component documents are available on the CSBS framework page, including the Core Examiner Guide, the Examiner Work Program and the Risk Tiering Worksheet. The CSBS press release announced the public release on September 16, 2026. The revised interagency model risk guidance is published by the Federal Reserve as SR 26-2. Coverage of the release and of the federal scope exclusion, along with the Wolters Kluwer readiness survey, appears in National Mortgage News, and the NYDFS risk assessment guidance is summarized by Sheppard Mullin.
Tags: #mortgageservicing #ai