← Back to Our Work
AI Evaluation Risk Taxonomy Strategy Consulting AI Governance

SORTIE: An AI Evaluator Taxonomy for Risk Measurement

A client building an AI governance and compliance platform (confidential) had real infrastructure already in production — monitoring dashboards, audit logs, evaluator prototypes — but no principled way to say what actually needed to be measured, for which use case, and why. We designed SORTIE — the evaluator taxonomy and risk matrix that turned ad-hoc monitoring into a structured measurement system.

Use case by evaluation branch risk matrix
Domain
AI Governance & Compliance
Scope
Strategy + Taxonomy Design
Deliverables
Taxonomy · Registry · Scenario Cards

The Problem

Most teams' idea of "AI risk" is one or two things they already know to worry about — usually hallucination, sometimes bias. That's not because those are the only risks that matter. It's because without a structured way to think about it, a team can only check for the risks they've already imagined. An AI output actually has many more dimensions that can go wrong than that, and most of them stay invisible until something breaks in production.

The client was building a platform for exactly this problem — helping other companies monitor and govern their AI systems — but their own internal evaluation approach had the same gap. Evaluators were named and categorized inconsistently, with no principled structure connecting them, and no clear answer to "for this specific AI use case, which risks actually apply, and how do we measure each one?"

What We Built: The Matrix, Not a Checklist

The core insight is that AI risk isn't a list — it's a matrix. Every real use case an AI system is deployed for gets crossed with every evaluation dimension that could apply to it. A customer support chatbot and a financial document parser don't share the same risk profile, and neither should be evaluated against the same short checklist.

The matrix below shows how required evaluation coverage differs by use case — nothing gets checked by accident, and nothing gets missed by default.

We designed a six-branch evaluator taxonomy covering the full surface of AI output evaluation, structured around a simple principle: each branch answers a distinct question a risk or compliance team would actually ask.

That question-first structure — rather than organizing by ML technique or tool category — is what makes the taxonomy usable by a compliance team, not just an engineering team. The six branches spell SORTIE — Safety, Output Quality, Robustness, Transparency, Integrity, Efficiency — which is also just what running one is: a targeted pass over the system to check every angle, not a scan for the one risk you already suspect.

SORTIE taxonomy tree — six branches with their evaluator subcategories

Each branch splits into specific evaluators — the full taxonomy runs 20+ deep, shown here at the top two levels.

Why a Taxonomy Alone Isn't Enough

Knowing that six branches of risk exist doesn't tell a team which evaluators to actually run, in what priority, or how to measure them cheaply. We built a registry specifying, for each evaluator, exactly how it should be measured — a deterministic check that runs on every output for free, versus an expensive judgment call that should be reserved for the cases that actually need it.

That distinction is where most teams overspend without realizing it: defaulting to an expensive AI-judged evaluation for something that could have been a simple, free, deterministic check. Getting that distinction right is often the single biggest cost lever in an evaluation system.

Making It Usable: Scenario Cards

The last piece connects the taxonomy to real use cases. For each scenario — a support chatbot, a document parser, a Q&A assistant, a content summarizer — we mapped exactly which evaluators are required, which are recommended, and which only apply conditionally. That mapping is what turns a taxonomy from a reference document into something a team actually uses every time they add a new AI use case, instead of improvising a custom evaluation plan from scratch each time.

Outcomes

  • A six-branch evaluator taxonomy covering 20+ distinct evaluators, mapped to the questions a risk or compliance team actually asks
  • A structured registry specifying measurement method, cost profile, and ownership for every evaluator — making the cheap-vs-expensive tradeoff explicit instead of defaulting to the expensive option everywhere
  • Scenario cards for the platform's highest-frequency use cases, turning "what do we need to check for this AI feature" into a repeatable process instead of a one-off decision
  • A rewritten strategy document connecting the taxonomy to the client's product roadmap, so the measurement backbone and the business plan point the same direction

What This Means for Any AI Product — Not Just Regulated Platforms

This pattern isn't only relevant to enterprise governance platforms. Any team shipping an AI feature is implicitly making a bet about what could go wrong, whether they've written it down or not. The matrix — real use cases crossed with the dimensions that actually matter for each — is the same tool whether you're a two-person startup evaluating one chatbot or a platform monitoring AI for dozens of enterprise customers. The size of the matrix changes. The need for one doesn't.

If your team is relying on a mental checklist instead of a structured way to know what to measure, that gap tends to stay invisible until something goes wrong in front of a customer. Read more on how we think about evaluation in how many AI agents your product actually needs and building an AI evaluation strategy, or explore AI Strategy Consulting.

Not sure what your AI system's actual risk surface looks like? Explore AI Strategy Consulting →

Not Sure What Your AI Risk Surface Looks Like?

30 minutes with a senior AI consultant. We'll help you map the real use case × risk matrix for your product, not just the checklist you already know.

Got Questions?

Send Us a Message

We'll reply within one business day.

+1 916 936 1544
Sacramento, CA
Waking up the assistant…