The AI Use-Case Priority Matrix
An AI use-case workshop can quickly turn into a feature-idea contest. Executives choose visible projects, developers choose easy projects, and operators choose whatever is most annoying today. Without common criteria, the loudest person sets the roadmap.
A strong AI use case is not the most impressive idea. It is work with material business value, repeatable evaluation, tolerable failure, and users who are willing to change how they work.
This matrix is a starting point for comparing five to ten candidates in a 60–90 minute workshop. The weights are not a universal standard. Adjust them to the organization’s industry, risk, and strategy.
Step 1: Rewrite the Idea as a Task
“Use AI in sales” and “automate support” cannot be scored. Use this grammar:
[Who] receives [what input], produces or performs [what output or action], and [who] verifies success using [what criteria].
For example: “A support agent receives the order and return policy plus conversation history, drafts a response, and a human agent approves policy compliance, accuracy, and resolution.”
Attach the current baseline: monthly volume, average handling and waiting time, error and rework rates, cost, and user satisfaction.
Step 2: Apply Hard Stops Before Scoring
Pause or redesign the task if any statement is true:
- There is no workflow owner or budget owner.
- No person or data source can determine success.
- Required data or system access cannot be approved.
- Value requires irreversible, high-impact action without human approval.
- Volume is too low to repay evaluation and operational investment.
- The real problem is a broken policy, dataset, or process rather than an AI problem.
AI applied to a broken process spreads the failure faster.
Step 3: Use the 100-Point Matrix
Score each criterion from one to five. Weighted points equal score ÷ 5 × weight.
| Criterion | Weight | 1 Point | 3 Points | 5 Points | Evidence |
|---|---|---|---|---|---|
| Business impact | 25 | Minor and indirect | Improves one team’s time or quality | Improves a core revenue, cost, or risk KPI | Baseline and financial assumptions |
| Technical feasibility | 20 | Data and integration unavailable | Some manual work and integration needed | Data, APIs, and permissions ready | Architecture and sample data |
| Evaluability | 15 | No clear answer or judgment rule | Human judgment possible | Reproducible tasks plus automated and human graders | Eval set and grading rules |
| Frequency and scale | 10 | Rare exception | Weekly or one team | Daily and repeated across teams | Volume and user count |
| User desirability | 10 | Resistance or added burden | Neutral, training needed | Clear pain and adoption intent | Interviews and observation |
| Risk fit | 15 | High-impact, irreversible, sensitive | Mitigated by controls and approval | Lower-impact, reversible, non-sensitive | Risk register and approval path |
| Operational readiness | 5 | No owner or support | Temporary owner | Operator, budget, and runbook plan | RACI and budget |
A five in risk fit does not mean “no AI risk.” It means the impact is limited, actions can be reversed, and least privilege, approval, and logging can control the exposure. The NIST AI RMF likewise treats risk as something mapped, measured, and managed in context.
Step 4: Read the Total and the Quadrant Together
A single total can allow high value to cancel out high risk. Put technical feasibility plus evaluability on the horizontal axis, business impact on the vertical axis, and show risk as a separate color or column.
| Quadrant | Interpretation | Action |
|---|---|---|
| High value, high executability | Start now | Run a PoC with real failures in the eval set |
| High value, low executability | Build the foundation first | Improve data, permissions, and process, then rescore |
| Low value, high executability | Learning use case | Limit to low-cost platform or training work |
| Low value, low executability | Stop | Remove from the active idea list |
High-impact irreversible work needs separate executive approval regardless of its total score.
Illustrative Scores
The following numbers are hypothetical.
| Candidate | Impact | Feasibility | Eval | Frequency | User | Risk | Ops | Total | Decision |
|---|---|---|---|---|---|---|---|---|---|
| Support response draft | 4 | 4 | 5 | 5 | 4 | 4 | 3 | 84 | Prioritize PoC |
| Automatic contract approval | 5 | 2 | 2 | 3 | 3 | 1 | 2 | 51 | Redesign the task |
| Meeting summary | 2 | 5 | 4 | 4 | 4 | 4 | 4 | 72 | Learning or self-service |
| Executive strategy advice | 4 | 2 | 1 | 1 | 3 | 2 | 1 | 45 | Narrow or pause |
If “automatic contract approval” becomes “contract summary and clause extraction,” risk and evaluability change. The purpose of the matrix is not to kill ideas. It is to cut them into safe, testable units.
A 60–90 Minute Workshop
- The workflow team explains each candidate and baseline in five minutes.
- Business, IT, security, and operations score independently.
- Discuss criteria where scores differ by two points or more.
- Attach required evidence and an owner to every disputed claim.
- Assign eval tasks, a PoC owner, and an end date to the top two.
- Record why the others are waiting and when they will be reviewed.
Because agent behavior can vary between runs, evaluation needs tasks, trials, graders, and traces rather than one demo. Anthropic’s evaluation guidance is a useful reference for designing that next step.
Blank Scorecard
| Candidate Task | Impact 25 | Feasibility 20 | Eval 15 | Frequency 10 | User 10 | Risk 15 | Ops 5 | Total | Hard Stop | Next Evidence and Owner |
|---|---|---|---|---|---|---|---|---|---|---|
The highest score is not automatically the first project. Choose the candidate that can be evaluated, can fail safely, and has someone prepared to own its operation.