What a rubric is
A criterion is one binary check on the deck. A judge — a person or an LLM — reads your criterion, looks at the deck, and answers pass or fail. Nothing else is available to them.
A set of criteria does two jobs at once. It measures difficulty, as the pass rate across the set. And it captures end-user happiness: would someone in your seat accept this deck. Mostly that means design, formatting, layout and structure — but a critical number stated wrong is an end-user problem too, and worth catching.
You write the criteria for your own task. You know why the third quadrant sits where it does and which segment belongs first — you built the task. Hand it to someone else, even someone in the same industry, and that context is gone.
Everything in this guide follows from one fact about the judge.
- The deck
- Your criterion, on its own
- The prompt
- Your input files — the CSV, the PDF
- The template
- Your other criteria
Write every criterion as if it is the only sentence the judge will ever read about this task. It is.
The four questions
Don't stare at the outputs waiting for criteria to arrive. Work these four questions in order — each one produces a different kind of criterion.
Questions 1 to 3 come before you open a single output. Question 4 comes after. Section 4 explains why that order is not optional.
The tags on the right are the two categories every criterion you write gets filed under. Question 1 produces one of them; the other three produce the other. Section 3 is the line between them.
Outcome and Quality
You write two kinds of criteria. A third kind runs on its own.
- Outcome
- The prompt put it there.
- Quality
- You'd have done it anyway.
That is the whole test, and it is one question: is this related to the prompt, or isn't it? Anything the prompt puts in scope is Outcome — including the input files you supplied with it. Everything you bring because you know the work is Quality.
Seen on one slide
The prompt was four words: Visualise the value chain. Four criteria come off the slide it produced — two of each.
- 1QualityThe title states the takeaway, not Value chain.Nobody asked for it. Everyone in your seat would write it that way.
- 2QualityIt is a chevron, reading left to right.A pie chart would also have been "a visual". The shape is judgment.
- 3OutcomeThere is a value-chain visual on the slide.The four words of the prompt, met. Explicit.
- 4OutcomeEvery stage in your input pack appears, and nothing else.Implicit: the files you supplied put those six stages in scope.
What the prompt asked for
Explicit asks are the ones written down: include a bridge, exactly 8 content slides.
Implicit asks are everything else related to the prompt — including the input files you supplied with it. If your PDF flagged growth in one segment, you raised that segment, so the deck reflecting it is still an Outcome ask even though the prompt never spelled it out.
You don't need to test everything in the prompt. Outcome criteria mostly pass. Test the ones that don't.
Everything requiring judgment
Verbosity and prose quality are always Quality, on every task. So is everything else that depends on knowing the work:
- Design — the template specifies grey and the output rendered it blue; the chart runs past the margin.
- Format — notes sit above sources, in italic, the way your industry does it.
- Layout — chart on the left, commentary on the right, so the slide reads in the order it's meant to.
- Narrative and prose — bullets that give the driver behind a number rather than restating it; titles pitched for the audience the deck is going to.
- Domain convention — FDD bullets don't read like consulting bullets, which don't read like banking bullets. In consulting they run long: full sentences, self-contained enough to read without anyone presenting them.
- Consistency — one colour per category across the deck; one naming convention per set of headers; one taxonomy throughout.
You write nothing
Runs automatically in the background on every deck. It covers unnecessary white space, overlapping shapes, overlapping text, font size and template match — and the mechanical annoyances underneath, like a chart rendered as a stack of drawn objects rather than a real chart. It also includes a gate: a file that won't open, or that triggers PowerPoint's repair prompt, scores zero. Structural still moves your overall pass rate, so don't be surprised by it.
Content checks are legitimate and a few are wanted — a critical figure, a segment that should be there, an explanation the slide can't do without. But they're the minority. The bulk of your set should be design, prose, verbosity and layout.
New writers get the boundary wrong more often than anything else. If answering the question needs judgment, it is Quality — even when every expert would agree. Unanimity doesn't make something Outcome.
| The prompt says | OutcomeWhat it asked for | QualityWhat judgment adds |
|---|---|---|
| "Visualise the value chain." | A value-chain visual is on the slide, carrying every stage in the input pack. | It is a chevron, reading left to right, one element per stage. |
| "Add a company overview slide." | The deck contains a company overview slide. | Business description top left, revenue split by segment on the right — the 2×2 anyone in your seat would build. |
| "Five slides for the ExCo." | The deck contains exactly five content slides. | Each title states a takeaway, in language pitched at an ExCo. |
Read the middle column on its own and notice how little of a good slide the prompt actually buys. Everything that makes it worth sending is in the third.
Unrequested content fails Outcome
The prompt asked for EBITDA and margin; the output added a revenue row. That's a fail. One word catches it cleanly:
Passes even when the output adds three rows nobody asked for.
Where criteria come from
Expectations first
Before you open a single output, write down what you'd expect on the slide. The useful framing: if five people in your seat built this slide, what would all five put on it?
- Company overview → a 2×2, business description top left, revenue split on the right.
- Revenue history for a business you know → the COVID dip visibly flagged, whether or not anyone asked.
- Competitive landscape quadrant → competitors segmented, highlighted by colour, logos present.
- Any segmented view → the key target named first, and any change in the segmentation explained where it happens. A deck that shows six segments on one slide and five on the next, without a word about the sixth, is a deck someone will send back.
Judge the slide as a whole as well as part by part. A slide can pass on every element and still fail to make its point.
It makes no difference whether your task is built on an old case with a finished deck behind it, or invented from scratch. Either works. What you're describing is what the slide should contain, not what some existing version happened to do.
Then error-hunt
Now open the ten outputs and log what's wrong.
The platform can draft criteria for you. Use it after you've written your own, not instead — the point of writing them cold is that it forces you to decide what good looks like before an output tells you.
Why that order
Start from the outputs and you end up describing them. You get a long list of small faults, most of them particular to the ten decks in front of you, and the criteria that actually define a good slide never get written — because nothing in front of you prompted them.
Your expectations don't move when the outputs do. Write those first and the error-hunt becomes a second pass that adds to a set, rather than the whole set.
Write the criterion when the same error turns up in deck after deck across your run set. One or two outputs getting something wrong is noise — that's one model having a bad run, not a failure mode. Eight of ten is the working guide, but the substance is recurrence, not the count.
Your task type changes the difficulty
From scratch, or an empty template. Expectation criteria are easy to write and they fail often — the unprompted pie chart is the classic. Layout is worth criteria here: where the chart sits, where the commentary sits, what reads first. When the template already places the chart, layout matters much less and judgment matters more.
In-place edits — add China alongside Great Britain and the US. Harder, because the agent mostly just replicates what's there. Still workable: the slide is divided into three verticals; the three charts are identical in layout, design and font.
- The title overrunning the subtitle and the rule, the clipped labels in the blue column, the collided figures in the headers — all Structural. Already caught, don't write them.
- The panel inconsistency and the duplicated flag — task-specific, invisible to a universal check, and yours to write.
The ten rules
One standard sits behind all ten. If five people read your criterion against the same deck — one of them an LLM — all five reach the same verdict. A criterion judges disagree on is worth nothing.
A good criterion passes three tests: it bears on whether the slide communicates to its audience, it targets one meaningful judgment, and judges land on the same answer every time. The ten rules below are how you get there.
Self-contained
The judge has the deck. Not your CSV, not your prompt, not your template.
No room for interpretation
There is a definition of "consulting-like". There is no agreement on it.
Name the thing, not the vibe
Everyone knows the AI-generated look when they see it. Nobody defines it the same way. Describe what's actually on the slide.
One judgment per criterion
Stacking leaves the judge guessing whether one hit is a pass, and it penalises an output that got half of it right. Split them.
The footnote sits above the source line.
Genuinely paired items are the exception — EBITDA and EBITDA margin can share one criterion.
Don't overfit
You are describing what good looks like, not the one output in front of you. A two-line title can be perfect.
Say what should be true
LLM judges read "does not" and "should not" badly, and a criterion built on absence is harder to verify than one built on presence. State the positive. Keep "avoids" for the cases where absence genuinely is the point.
No "X rather than Y"
If the bars are blue, they're blue. The contrastive tail adds nothing and gives judges a second thing to disagree about.
Name it, don't point at it
A criterion is read alone, with nothing before it. Relational references have nothing to point at.
One data point is enough
You're checking whether the model carried the numbers across correctly, not auditing the deck. Don't make a judge verify a series.
Keep it short
If you have to read your own criterion twice, so will everyone else — and they'll each land somewhere different.
What not to write
Anything already automated
Universal checks run on every deck without you. Font size, template match, collisions, overlaps. Writing them again double-penalises the same mistake.
The AI "dashboard" look — big coloured cards, oversized metric boxes — is covered centrally too. It's a rule everywhere, on every deck, so you don't need a criterion for it.
The list above isn't everything the universal checks cover. Run them on your own outputs before you start writing — what they flag is your working list of what not to duplicate. Anything borderline, ask in the channel.
Anything generic about the design
Legibility, contrast, overlapping elements, too much white space, "this just looks wrong" — none of it is yours. It's caught centrally and writing it again double-penalises the deck.
The exception is narrow: a case so specific that no universal check could know about it. The brand guide that names a colour, the prompt that specified which shape sits on top, the footnote convention your industry insists on. If you could copy the criterion into another task unchanged, it belongs to the universal checks, not to you.
But do write the task-specific version
The universal library catches what's true of every deck. It cannot catch what's true of yours:
- The prompt asked for the ellipses on top of the arrows; the output reversed them.
- In your industry, notes sit above sources, in italic.
- The brand guide specifies grey; the output rendered it blue.
- The template's horizontal rule is missing on half the slides.
Anything experts disagree on
A criterion has to hold at least within your own domain. When it doesn't, you're teaching the model one house style and calling it a standard.
When you're not 100% sure, leave it out. A smaller set of criteria everyone agrees on beats a longer one that judges argue over.
One-off errors
If it doesn't recur across decks, it isn't a failure mode. One bad output is one bad output.
Calibrating the set
How many
Aim for 20 criteria per deck — not per slide. More only if the task genuinely needs it. Don't pad the set to hit a number: a task with 40 weak criteria is worse than one with 15 strong ones.
Four to five pages of task is the sweet spot. Fewer and the model finds it too easy. Longer is fine when the work is simple and repetitive — a rebrand plus a numbers refresh across an existing deck — because it isn't asking for many different things at once.
Tag each one
Outcome or Quality. Deck-wide or slide-specific. Deck-wide means it applies to every slide — the title sits above the blue rule and the subtitle below it.
Make it hard
A task the models sail through adds nothing to the dataset. Outcome criteria will mostly pass — that's expected, and inventing stricter ones to force failures is the wrong move. Quality is the half you can move. If your set is passing too easily, the headroom is in the judgment criteria.
Criteria quality first, difficulty second
Get the criteria right, then look at how the set performs. A hard-looking pass rate built on weak criteria tells nobody anything.
Keep your failure modes diverse
If everything your set catches traces back to one root cause, a single upstream fix takes the whole task from hard to trivial. Spread the failures across different kinds of mistake.
No weights
Criteria are not weighted. Every one counts the same, so a criterion you're lukewarm about dilutes the ones you care about. That's the argument for cutting rather than keeping.
Before you submit
Run the set against this. Anything you can't tick, fix or cut.
- Every criterion is readable with the deck alone — no reference to the prompt, the CSV or the template.
- No criterion tests two unrelated things.
- Phrased positively — "avoids" only where absence is the point. No "rather than". No "this" or "that".
- Nothing here is already covered by the automated checks.
- Nothing here is something I'm less than certain about.
- Every error-derived criterion describes something that recurs deck after deck, not a one-off.
- Each one is tagged Outcome or Quality, deck-wide or slide-specific.
- Read cold: would someone else grade this deck exactly as I would?
Worked examples
Real criteria from a finished Atlas set — a thirteen-slide commercial due diligence deck on First Mile, a UK commercial waste company. Each one shows what the rules look like in practice.
How a set is put together
Every criterion carries three things: an id, a category — Outcome or Quality — and a slide, where it applies to one slide rather than the whole deck. That's all the structure there is.
Outcome
The prompt asked for eight. The clause that earns its place is the second one — without it, judges disagree about whether the cover counts, and the criterion produces a different verdict depending on who reads it.
The "only" device from section 3, in the wild — it fails an output that helpfully adds a revenue row. Also the legitimate use of stacking: EBITDA and its margin are one judgment, not two.
Enumerated rather than left as "all stages of the value chain". The judge has no prompt and no domain brief to resolve "all" against — so the criterion names them.
Asked for in the prompt, so Outcome. "Or in close proximity" is deliberate slack — without it the criterion fails a perfectly readable chart on a few pixels, which is overfitting by another route.
One number, with a tolerance. Rule 9 in practice: you're checking whether the model can compute and place a CAGR, not auditing seven years of revenue. The range stops a judge failing an output over a rounding convention.
Self-containment done well. The judge can't check the figures against the source pack, but internal consistency is visible on the slide — and mixing reported with pre-exceptional EBITDA across two adjacent years is exactly the error that would get the deck sent back.
Quality — deck-wide
"The titles are action titles" made testable. Nobody asked for it in the prompt, everyone in consulting expects it, and every clause here exists to close a gap a judge would otherwise fill differently.
A numbers check that needs no source file at all. Margin against revenue and EBITDA, growth against the two years it spans — all of it is on the slide, so the judge can do the arithmetic.
The general rule, then one concrete instance of it. The example after the colon costs eight words and removes any argument about what alignment means here.
Error-derived, and it recurs deck after deck — models re-pick colours per slide. Note it doesn't prescribe which colour: prescribing one would overfit to a single output.
The right use of "avoids" — here absence genuinely is the point, so there's no positive form to state instead. Nine words. Rule 10 says that's a feature.
Quality — the general and the specific
The general form. It travels to any deck, which is what makes it worth writing once at deck level rather than per slide.
The specific form of the one above, and the pair is instructive: the general criterion catches whether the model highlights anything, the specific one catches whether it highlights the right years in every exhibit. The second sentence closes the loophole where the model distinguishes FY20 by making every other year a different colour too.
Internal consistency rather than a house style. It doesn't say which convention to use, so it can't teach the model one person's habit — it only requires that the deck pick one and hold it.
The deck that segments the market six ways on one slide and five ways on the next, saying nothing about the sixth. A reader notices in a second; the model never does. Note that it doesn't forbid the change — it requires the deck to own it.
And one that didn't make it
Legend placement is a preference, not a standard — two people with the same background will place them differently and both be right. A criterion built on a preference teaches the model one person's habit. When you're not sure, leave it out.
The same set shown end to end: the First Mile prompt, the expectations written before any output existed, the ten runs, and which criteria came from which.