Enterprise / high-volume plans available on request.
Anonymous users get 10 free calls/day without a key.
Free resource
SR 11-7 GenAI Documentation Checklist
A practical checklist for documenting an LLM or GenAI tool using the SR 11-7 model-risk discipline —
what a validator or examiner actually expects to see, not a summary of the regulation. Enter your work
email to unlock it; it prints cleanly to PDF from your browser.
Where this stands as of September 2026: SR 26-2 (issued April 17, 2026)
replaced the original SR 11-7 guidance — and, like SR 11-7 before it, explicitly excludes generative and
agentic AI from its formal model-risk validation scope, with a dedicated interagency framework for AI still to
come. SR 26-2 itself is mainly aimed at banks over $30B in assets. In practice, this means most GenAI tools
aren't formally required to pass this exact test — but the SR 11-7 documentation discipline is still what
examiners, boards, and prudent institutions of any size reach for as the working standard while dedicated
AI guidance is pending, which is why this checklist still reflects it.
Enter your work email to unlock the full checklist
Unlock above to view all 24 checklist items across 4 sections.
1 · Model Development & Documentation
Tool inventory entry — name, owner, business purpose, and every finance calculation or workflow the tool performs.
Model/vendor description — base LLM, version, any fine-tuning or system prompt engineering applied, and MCP tools or plugins it can call.
Intended use statement — what the tool is approved for, and explicit exclusions (e.g. "not approved for client-facing investment advice").
Data inputs and outputs — what data the tool consumes, where it comes from, and what format/precision its outputs are expected to have.
Known limitations — documented failure modes, edge cases, and calculation types the tool is not validated for.
Change log — version history of prompt, model, or tool changes, each with a date and rationale.
2 · Independent Validation
Ground-truth test suite — a set of test cases with analytically or independently computed correct answers, not model-judged.
Documented pass criteria — a defined tolerance (e.g. "within 0.01% of the closed-form value") agreed before testing, not chosen after seeing results.
Validator independence — testing performed by someone who did not build the tool, per SR 11-7's effective challenge requirement.
Edge-case and stress testing — behavior under malformed input, out-of-range values, and adversarial prompts, not just the happy path.
Reproducibility — can the same test suite be re-run by someone else and produce the same pass/fail result.
Written validation report — methodology, test results, and a explicit conclusion on whether the tool is fit for its intended use.
3 · Ongoing Monitoring
Re-validation trigger policy — what changes (model upgrade, prompt edit, new tool call) require re-running the test suite before redeployment.
Production monitoring plan — how output quality is sampled and checked after deployment, not just at launch.
Drift detection — a process for noticing if the underlying model's behavior changes after a vendor-side update you don't control.
Incident log — a record of any known wrong outputs in production, their impact, and remediation.
Re-validation schedule — a fixed cadence (e.g. quarterly) for re-running the full test suite regardless of whether changes were made.
User feedback channel — a way for people using the tool's output to flag suspected errors back to the model owner.
4 · Governance & Accountability
Named model owner — a specific accountable person, not a team or department.
Risk tiering — the tool's model-risk tier based on materiality and reliance, consistent with your firm's existing model risk framework.
Sign-off record — evidence the tool was formally approved for its intended use by whoever holds that authority at your firm.
SR 26-2 alignment — if this is a GenAI tool outside your formal validation scope, documentation showing why, per the Fed's SR 26-2 guidance.
EU AI Act classification — if operating in-scope, a documented risk classification and, for high-risk uses, a traceability record.
Third-party/vendor disclosure — if using a vendor's MCP server or API, documentation of what you validated versus what you're relying on the vendor's own claims for.
Need this filled in, not just outlined?
We build the ground-truth test suite, run the validation, and write the report — sections 1 and 2 above, done for your specific tool.