QGI · documentation-readiness model · SR 26-2 issued 2026-04-17
SR 26-2 generative AI model risk management readiness score
SR 26-2, issued 17 April 2026 by the Fed, OCC and FDIC, replaced SR 11-7 and placed generative and agentic AI outside its scope. Banks still have to govern those systems, and examiners are already asking how. This self-assessment scores your documentation readiness across the ten controls that vendor analyses and FINRA guidance name most often: inventory, risk tiering, approval gates, hallucination testing, human review, output logging, kill switch, vendor review, determinism and an evidence pack. It is a readiness model, not supervisory advice.
- Inventory score
- 50/100
- Tiering score
- 0/100
- Controls score
- 50/100
- Testing score
- 30/100
- Monitoring score
- 25/100
- Vendor score
- 0/100
- Gap 1 (surfaces first)
- Use cases are not risk-tiered with a written proportionality rationale
- Gap 2 (surfaces first)
- Decisioning outputs are not reproducible run to run
- Gap 3 (surfaces first)
- No examiner-ready evidence package
- Benchmark (Harness, 2026-09-10)
- 44% have a real agent inventory; 33% can stop an agent; 19% have controls that catch failures
Score 32/100, tier Exposed. The first gap an examiner would find: use cases are not risk-tiered with a written proportionality rationale.
Assumptions and sources (8)
| Constant | Value | Basis |
|---|---|---|
| Scoring | 0 / 5 / 10 points per question; weighted sum / 110 x 100 | estimateStudio scoring model; weights are a documentation-readiness estimate, not a regulator scale |
| Weights | inventory 1.5, tiering 1.5, testing 1.5, approval gate 1.0, human review 1.0, logging 1.0, kill switch 1.0, determinism 1.0, evidence pack 1.0, vendor review 0.5 | estimateMapped to controls named in vendor SR 26-2 analyses (Domino) |
| Tiers | <40 Exposed, 40 to 69 Building, 70 to 89 Exam-ready, 90+ Leading | estimateStudio labels |
| Size adjustment | -5 points for >$100B assets with >20 live use cases | estimateLarger institutions face higher examiner expectation (studio estimate) |
| Control set: inventory incl. agent workflows and third-party LLMs | PiTech banking AI inventory guide | sourcedPiTech |
| Control set: hallucination / bias testing, sampling, escalation | FINRA 2026 Regulatory Oversight Report, GenAI section | sourcedFINRA |
| SR 26-2 scope | Generative and agentic AI out of scope; issued 2026-04-17 | sourcedFederal Reserve SR 26-2 |
| Benchmarks | 44% real inventory, 33% can stop an agent, 19% effective controls | sourcedHarness, State of AI Agent Controls, 2026-09-10 |
Key takeaways
- SR 26-2 replaced SR 11-7 on 2026-04-17 and left generative and agentic AI out of scope. Out of scope is not out of governance.
- Examiners ask three things first: is it in the inventory, can you stop it, and can you show the test that caught its failures.
- Only 44% of firms have a real agent inventory and 33% can actually shut an agent down (Harness, 2026-09-10).
- Deterministic decision layers are the part of a GenAI stack a model-risk team can validate like a model.
- This is a documentation-readiness score, not supervisory advice and not a prediction of exam findings.
How the SR 26-2 readiness score is calculated
Ten control questions, each scored none, partial or full. Each question carries a weight: inventory, tiering and testing count 1.5, most controls count 1.0, vendor review 0.5. The weighted sum is normalised to 100. Institutions over $100B in assets running more than twenty live use cases lose five points, because examiner expectation scales with footprint.
The control set is not invented. It follows the controls named in vendor readings of the letter, the FINRA 2026 GenAI guidance and the OCC's notice that an interagency request for information on generative and agentic AI is coming (OCC NR 2026-29). QGI builds the deterministic and evidence layers the score asks about; TheoSym's AI consulting practice runs the readiness work with your model-risk team.
- Inventory: models, agent workflows, third-party LLMs
- Tiering: materiality and downstream impact, with a written rationale
- Testing: repeatable eval suite with hallucination and bias thresholds
- Controls: approval gate, human review, kill switch with rollback
- Monitoring and evidence: logging, sampled review, escalation, examiner pack
SR 26-2 vs SR 11-7: what changed for generative AI
SR 11-7 was a checklist. SR 26-2 is a principle: proportionality, judged by the institution and defended to the examiner. Vendors summarised the shift as "SR 26-2 removed the checklist" (CIMCON). The letter explicitly places generative and agentic AI outside its definition of a model, calling them novel and rapidly evolving.
That leaves a gap, not a pass. The systems still fall under existing risk frameworks, third-party risk guidance and consumer protection rules. A bank that cannot show an inventory, a tiering rationale and test evidence for its GenAI use cases has a documentation problem the day an examiner asks. The AI Factory workproofs show how QGI packages that evidence.
What is a good SR 26-2 readiness score?
Seventy or above means every core control exists in documented form and the evidence pack is at least in progress. Below forty means an examiner conversation would surface gaps in the first ten minutes. Most institutions land in the middle: logging exists, testing is ad hoc, the kill switch has never been rehearsed.
The benchmark is unflattering. Across 700 engineering leaders, 77% believed they had a complete agent inventory and 44% did; 76% believed they could shut an agent down quickly and 33% could (Harness, 2026-09-10). Score honestly. The number is for you, not the examiner.
When do you need SR 26-2 readiness work?
Before the request for information lands, and before the next exam cycle. Buyer guides describe a focused 60-day scope: re-baseline the inventory, tier by impact, map controls to tiers, close the GenAI and agentic carve-out with explicit controls, package the evidence (PiTech buyer’s guide).
Where a use case makes or supports a decision a customer can challenge, credit, eligibility, pricing, reporting, the determinism question matters most. QGI, led by Dr. Sam Sammane, Founder & CEO, builds decision layers that return the same output for the same inputs on every run, with a versioned rule set and an audit trail. Announcements are in the press center.
SR 26-2 generative AI model risk management: questions people ask
Does SR 26-2 apply to generative AI and AI agents?+
No. SR 26-2, issued 17 April 2026 by the Fed, OCC and FDIC, replaces SR 11-7 but places generative and agentic AI outside its scope as novel and rapidly evolving. Banks must still govern those systems under existing risk frameworks, and examiners are already asking how, with an interagency request for information expected.
What do examiners ask about GenAI if it is out of scope?+
In practice: whether the AI inventory includes GenAI models, agent workflows and third-party LLMs; how use cases are risk-tiered; whether there is an approval gate, hallucination and bias testing, human review of material outputs, output logging with escalation, and whether an agent can be halted and rolled back.
How long does an SR 26-2 readiness effort take?+
Vendor guides describe a focused 60-day scope: re-baseline the model and AI inventory, tier by impact, map controls to tiers with a documented proportionality judgement, close the GenAI and agentic carve-out with explicit controls, and package examiner-ready evidence. The score highlights which of those steps is weakest.
Why does determinism matter for model risk in decisioning?+
Validation depends on reproducibility. A deterministic decision layer returns the same output for the same inputs on every run, which lets model-risk teams test, explain and document it like a traditional model. Generative outputs vary run to run, which is why regulators treat them as harder to validate.
Is this readiness score supervisory or legal advice?+
No. It is a documentation-readiness model built from public vendor analyses, FINRA guidance and the SR 26-2 text. It does not represent any regulator, predict exam findings or replace counsel. Use it to find the weakest control before an examiner does, then verify against the primary texts and your own policies.
Next step
Take the result to the QGI team
Send the score. Dr. Sam Sammane, QGI Founder & CEO, or an engineer who has shipped replies within one business day with the gaps that matter first.
Prefer to talk? +1 657-888-0688 or contact@theosym.com