A client building an AI governance and compliance platform (confidential) had real infrastructure already in production — monitoring dashboards, audit logs, evaluator prototypes — but no principled way to say what actually needed to be measured, for which use case, and why. We designed SORTIE — the evaluator taxonomy and risk matrix that turned ad-hoc monitoring into a structured measurement system.
Most teams' idea of "AI risk" is one or two things they already know to worry about — usually hallucination, sometimes bias. That's not because those are the only risks that matter. It's because without a structured way to think about it, a team can only check for the risks they've already imagined. An AI output actually has many more dimensions that can go wrong than that, and most of them stay invisible until something breaks in production.
The client was building a platform for exactly this problem — helping other companies monitor and govern their AI systems — but their own internal evaluation approach had the same gap. Evaluators were named and categorized inconsistently, with no principled structure connecting them, and no clear answer to "for this specific AI use case, which risks actually apply, and how do we measure each one?"
The core insight is that AI risk isn't a list — it's a matrix. Every real use case an AI system is deployed for gets crossed with every evaluation dimension that could apply to it. A customer support chatbot and a financial document parser don't share the same risk profile, and neither should be evaluated against the same short checklist.
The matrix below shows how required evaluation coverage differs by use case — nothing gets checked by accident, and nothing gets missed by default.
We designed a six-branch evaluator taxonomy covering the full surface of AI output evaluation, structured around a simple principle: each branch answers a distinct question a risk or compliance team would actually ask.
That question-first structure — rather than organizing by ML technique or tool category — is what makes the taxonomy usable by a compliance team, not just an engineering team. The six branches spell SORTIE — Safety, Output Quality, Robustness, Transparency, Integrity, Efficiency — which is also just what running one is: a targeted pass over the system to check every angle, not a scan for the one risk you already suspect.
Each branch splits into specific evaluators — the full taxonomy runs 20+ deep, shown here at the top two levels.
Knowing that six branches of risk exist doesn't tell a team which evaluators to actually run, in what priority, or how to measure them cheaply. We built a registry specifying, for each evaluator, exactly how it should be measured — a deterministic check that runs on every output for free, versus an expensive judgment call that should be reserved for the cases that actually need it.
That distinction is where most teams overspend without realizing it: defaulting to an expensive AI-judged evaluation for something that could have been a simple, free, deterministic check. Getting that distinction right is often the single biggest cost lever in an evaluation system.
The last piece connects the taxonomy to real use cases. For each scenario — a support chatbot, a document parser, a Q&A assistant, a content summarizer — we mapped exactly which evaluators are required, which are recommended, and which only apply conditionally. That mapping is what turns a taxonomy from a reference document into something a team actually uses every time they add a new AI use case, instead of improvising a custom evaluation plan from scratch each time.
This pattern isn't only relevant to enterprise governance platforms. Any team shipping an AI feature is implicitly making a bet about what could go wrong, whether they've written it down or not. The matrix — real use cases crossed with the dimensions that actually matter for each — is the same tool whether you're a two-person startup evaluating one chatbot or a platform monitoring AI for dozens of enterprise customers. The size of the matrix changes. The need for one doesn't.
If your team is relying on a mental checklist instead of a structured way to know what to measure, that gap tends to stay invisible until something goes wrong in front of a customer. Read more on how we think about evaluation in how many AI agents your product actually needs and building an AI evaluation strategy, or explore AI Strategy Consulting.
Not sure what your AI system's actual risk surface looks like? Explore AI Strategy Consulting →
We'll reply within one business day.