Project Atlas  ·  Contributor guide

Writing Atlas Rubrics

Atlas is a dataset of slide-building tasks — consulting, banking, FP&A, corporate, consumer — built to stay hard for the next 12 to 18 months. You write the criteria your task's outputs get graded against — this guide is how to write them well.

Read once. Then keep it open while you write. ~15 minutes. Draft — September 2026
Section 1

What a rubric is

A criterion is one binary check on the deck. A judge — a person or an LLM — reads your criterion, looks at the deck, and answers pass or fail. Nothing else is available to them.

A set of criteria does two jobs at once. It measures difficulty, as the pass rate across the set. And it captures end-user happiness: would someone in your seat accept this deck. Mostly that means design, formatting, layout and structure — but a critical number stated wrong is an end-user problem too, and worth catching.

You write the criteria for your own task. You know why the third quadrant sits where it does and which segment belongs first — you built the task. Hand it to someone else, even someone in the same industry, and that context is gone.

Everything in this guide follows from one fact about the judge.

The judge sees
  • The deck
  • Your criterion, on its own
The judge never sees
  • The prompt
  • Your input files — the CSV, the PDF
  • The template
  • Your other criteria

Write every criterion as if it is the only sentence the judge will ever read about this task. It is.

Section 2

The four questions

Don't stare at the outputs waiting for criteria to arrive. Work these four questions in order — each one produces a different kind of criterion.

1What does the prompt actually ask for?OutcomeExplicit asks, and anything the prompt puts in scope
2What would anyone in your domain do here without being told?QualityThe judgment layer
3What key question would your audience still have after seeing this slide?QualityContent and narrative, usually
4Looking at the outputs, what would you still do differently?QualityError-derived — Outcome if the prompt was missed

Questions 1 to 3 come before you open a single output. Question 4 comes after. Section 4 explains why that order is not optional.

The tags on the right are the two categories every criterion you write gets filed under. Question 1 produces one of them; the other three produce the other. Section 3 is the line between them.

Section 3

Outcome and Quality

You write two kinds of criteria. A third kind runs on its own.

Outcome
The prompt put it there.
Quality
You'd have done it anyway.

That is the whole test, and it is one question: is this related to the prompt, or isn't it? Anything the prompt puts in scope is Outcome — including the input files you supplied with it. Everything you bring because you know the work is Quality.

Seen on one slide

The prompt was four words: Visualise the value chain. Four criteria come off the slide it produced — two of each.

  1. 1QualityThe title states the takeaway, not Value chain.Nobody asked for it. Everyone in your seat would write it that way.
  2. 2QualityIt is a chevron, reading left to right.A pie chart would also have been "a visual". The shape is judgment.
  3. 3OutcomeThere is a value-chain visual on the slide.The four words of the prompt, met. Explicit.
  4. 4OutcomeEvery stage in your input pack appears, and nothing else.Implicit: the files you supplied put those six stages in scope.
Same slide, same four-word prompt. Outcome criteria mostly pass — that is expected. Quality is the half you can move.
Outcome

What the prompt asked for

Explicit asks are the ones written down: include a bridge, exactly 8 content slides.

Implicit asks are everything else related to the prompt — including the input files you supplied with it. If your PDF flagged growth in one segment, you raised that segment, so the deck reflecting it is still an Outcome ask even though the prompt never spelled it out.

You don't need to test everything in the prompt. Outcome criteria mostly pass. Test the ones that don't.

Quality

Everything requiring judgment

Verbosity and prose quality are always Quality, on every task. So is everything else that depends on knowing the work:

  • Design — the template specifies grey and the output rendered it blue; the chart runs past the margin.
  • Format — notes sit above sources, in italic, the way your industry does it.
  • Layout — chart on the left, commentary on the right, so the slide reads in the order it's meant to.
  • Narrative and prose — bullets that give the driver behind a number rather than restating it; titles pitched for the audience the deck is going to.
  • Domain convention — FDD bullets don't read like consulting bullets, which don't read like banking bullets. In consulting they run long: full sentences, self-contained enough to read without anyone presenting them.
  • Consistency — one colour per category across the deck; one naming convention per set of headers; one taxonomy throughout.
Structural

You write nothing

Runs automatically in the background on every deck. It covers unnecessary white space, overlapping shapes, overlapping text, font size and template match — and the mechanical annoyances underneath, like a chart rendered as a stack of drawn objects rather than a real chart. It also includes a gate: a file that won't open, or that triggers PowerPoint's repair prompt, scores zero. Structural still moves your overall pass rate, so don't be surprised by it.

Content checks are legitimate and a few are wanted — a critical figure, a segment that should be there, an explanation the slide can't do without. But they're the minority. The bulk of your set should be design, prose, verbosity and layout.

The trap

New writers get the boundary wrong more often than anything else. If answering the question needs judgment, it is Quality — even when every expert would agree. Unanimity doesn't make something Outcome.

The prompt says OutcomeWhat it asked for QualityWhat judgment adds
"Visualise the value chain."A value-chain visual is on the slide, carrying every stage in the input pack.It is a chevron, reading left to right, one element per stage.
"Add a company overview slide."The deck contains a company overview slide.Business description top left, revenue split by segment on the right — the 2×2 anyone in your seat would build.
"Five slides for the ExCo."The deck contains exactly five content slides.Each title states a takeaway, in language pitched at an ExCo.

Read the middle column on its own and notice how little of a good slide the prompt actually buys. Everything that makes it worth sending is in the third.

Unrequested content fails Outcome

The prompt asked for EBITDA and margin; the output added a revenue row. That's a fail. One word catches it cleanly:

Doesn't catch it
The table shows EBITDA and EBITDA margin.

Passes even when the output adds three rows nobody asked for.

Catches it
The table shows only EBITDA and EBITDA margin.
Section 4

Where criteria come from

Expectations first

Before you open a single output, write down what you'd expect on the slide. The useful framing: if five people in your seat built this slide, what would all five put on it?

  • Company overview → a 2×2, business description top left, revenue split on the right.
  • Revenue history for a business you know → the COVID dip visibly flagged, whether or not anyone asked.
  • Competitive landscape quadrant → competitors segmented, highlighted by colour, logos present.
  • Any segmented view → the key target named first, and any change in the segmentation explained where it happens. A deck that shows six segments on one slide and five on the next, without a word about the sixth, is a deck someone will send back.

Judge the slide as a whole as well as part by part. A slide can pass on every element and still fail to make its point.

It makes no difference whether your task is built on an old case with a finished deck behind it, or invented from scratch. Either works. What you're describing is what the slide should contain, not what some existing version happened to do.

Then error-hunt

Now open the ten outputs and log what's wrong.

The platform can draft criteria for you. Use it after you've written your own, not instead — the point of writing them cold is that it forces you to decide what good looks like before an output tells you.

Why that order

Start from the outputs and you end up describing them. You get a long list of small faults, most of them particular to the ten decks in front of you, and the criteria that actually define a good slide never get written — because nothing in front of you prompted them.

Your expectations don't move when the outputs do. Write those first and the error-hunt becomes a second pass that adds to a set, rather than the whole set.

Does it recur across decks?

Write the criterion when the same error turns up in deck after deck across your run set. One or two outputs getting something wrong is noise — that's one model having a bad run, not a failure mode. Eight of ten is the working guide, but the substance is recurrence, not the count.

Your task type changes the difficulty

From scratch, or an empty template. Expectation criteria are easy to write and they fail often — the unprompted pie chart is the classic. Layout is worth criteria here: where the chart sits, where the commentary sits, what reads first. When the template already places the chart, layout matters much less and judgment matters more.

In-place editsadd China alongside Great Britain and the US. Harder, because the agent mostly just replicates what's there. Still workable: the slide is divided into three verticals; the three charts are identical in layout, design and font.

A three-panel market spend slide showing UK, US and China columns, with an overrunning title, clipped labels in the right-hand column, and inconsistent row ordering between panels.
A real output — UK, US and China The third panel was the ask. What a writer catches here is the inconsistency between panels: the rows run in a different order in each one, and the China panel carries two flags. Those are the criteria to write — the panels are identical in layout, design and font, and each carries its own country's flag.
  • The title overrunning the subtitle and the rule, the clipped labels in the blue column, the collided figures in the headers — all Structural. Already caught, don't write them.
  • The panel inconsistency and the duplicated flag — task-specific, invisible to a universal check, and yours to write.
Section 5

The ten rules

One standard sits behind all ten. If five people read your criterion against the same deck — one of them an LLM — all five reach the same verdict. A criterion judges disagree on is worth nothing.

A good criterion passes three tests: it bears on whether the slide communicates to its audience, it targets one meaningful judgment, and judges land on the same answer every time. The ten rules below are how you get there.

1

Self-contained

The judge has the deck. Not your CSV, not your prompt, not your template.

Don't write
The FY24 revenue figure matches the source CSV.
Write
The FY24 revenue figure is 1,234.
2

No room for interpretation

There is a definition of "consulting-like". There is no agreement on it.

Don't write
The slide titles are consulting-like.
Write
Each slide title states a takeaway: a conclusion about the slide's content, expressed as a full clause.
3

Name the thing, not the vibe

Everyone knows the AI-generated look when they see it. Nobody defines it the same way. Describe what's actually on the slide.

Don't write
The slide avoids AI-looking boxes.
Write
The figures in the commentary box are set at or below the font size of the slide body text.
4

One judgment per criterion

Stacking leaves the judge guessing whether one hit is a pass, and it penalises an output that got half of it right. Split them.

Don't write
The title is an action title and the footnote sits above the source line.
Write — as two criteria
The title states a takeaway.
The footnote sits above the source line.

Genuinely paired items are the exception — EBITDA and EBITDA margin can share one criterion.

5

Don't overfit

You are describing what good looks like, not the one output in front of you. A two-line title can be perfect.

Don't write
The title is a one-line takeaway.
Write
The title is a takeaway.
6

Say what should be true

LLM judges read "does not" and "should not" badly, and a criterion built on absence is harder to verify than one built on presence. State the positive. Keep "avoids" for the cases where absence genuinely is the point.

Don't write
The commentary bullets do not just repeat the figures shown in the chart.
Write
Each commentary bullet gives a driver or a trend behind the figures.
7

No "X rather than Y"

If the bars are blue, they're blue. The contrastive tail adds nothing and gives judges a second thing to disagree about.

Don't write
The bars are blue rather than green.
Write
The bars are blue.
8

Name it, don't point at it

A criterion is read alone, with nothing before it. Relational references have nothing to point at.

Don't write
This chart uses the same colour scheme as the one above.
Write
The revenue chart and the margin chart use the same colour scheme.
9

One data point is enough

You're checking whether the model carried the numbers across correctly, not auditing the deck. Don't make a judge verify a series.

Don't write
All revenue figures from FY19 to FY24 are correct.
Write
The FY24 revenue figure is 1,234.
10

Keep it short

If you have to read your own criterion twice, so will everyone else — and they'll each land somewhere different.

Section 6

What not to write

Anything already automated

Universal checks run on every deck without you. Font size, template match, collisions, overlaps. Writing them again double-penalises the same mistake.

The AI "dashboard" look — big coloured cards, oversized metric boxes — is covered centrally too. It's a rule everywhere, on every deck, so you don't need a criterion for it.

The full list

The list above isn't everything the universal checks cover. Run them on your own outputs before you start writing — what they flag is your working list of what not to duplicate. Anything borderline, ask in the channel.

Anything generic about the design

Legibility, contrast, overlapping elements, too much white space, "this just looks wrong" — none of it is yours. It's caught centrally and writing it again double-penalises the deck.

The exception is narrow: a case so specific that no universal check could know about it. The brand guide that names a colour, the prompt that specified which shape sits on top, the footnote convention your industry insists on. If you could copy the criterion into another task unchanged, it belongs to the universal checks, not to you.

But do write the task-specific version

The universal library catches what's true of every deck. It cannot catch what's true of yours:

  • The prompt asked for the ellipses on top of the arrows; the output reversed them.
  • In your industry, notes sit above sources, in italic.
  • The brand guide specifies grey; the output rendered it blue.
  • The template's horizontal rule is missing on half the slides.

Anything experts disagree on

A criterion has to hold at least within your own domain. When it doesn't, you're teaching the model one house style and calling it a standard.

The test

When you're not 100% sure, leave it out. A smaller set of criteria everyone agrees on beats a longer one that judges argue over.

One-off errors

If it doesn't recur across decks, it isn't a failure mode. One bad output is one bad output.

Section 7

Calibrating the set

How many

Aim for 20 criteria per deck — not per slide. More only if the task genuinely needs it. Don't pad the set to hit a number: a task with 40 weak criteria is worse than one with 15 strong ones.

Four to five pages of task is the sweet spot. Fewer and the model finds it too easy. Longer is fine when the work is simple and repetitive — a rebrand plus a numbers refresh across an existing deck — because it isn't asking for many different things at once.

Tag each one

Outcome or Quality. Deck-wide or slide-specific. Deck-wide means it applies to every slide — the title sits above the blue rule and the subtitle below it.

Make it hard

A task the models sail through adds nothing to the dataset. Outcome criteria will mostly pass — that's expected, and inventing stricter ones to force failures is the wrong move. Quality is the half you can move. If your set is passing too easily, the headroom is in the judgment criteria.

Criteria quality first, difficulty second

Get the criteria right, then look at how the set performs. A hard-looking pass rate built on weak criteria tells nobody anything.

Keep your failure modes diverse

If everything your set catches traces back to one root cause, a single upstream fix takes the whole task from hard to trivial. Spread the failures across different kinds of mistake.

No weights

Criteria are not weighted. Every one counts the same, so a criterion you're lukewarm about dilutes the ones you care about. That's the argument for cutting rather than keeping.

Section 8

Before you submit

Run the set against this. Anything you can't tick, fix or cut.

  • Every criterion is readable with the deck alone — no reference to the prompt, the CSV or the template.
  • No criterion tests two unrelated things.
  • Phrased positively — "avoids" only where absence is the point. No "rather than". No "this" or "that".
  • Nothing here is already covered by the automated checks.
  • Nothing here is something I'm less than certain about.
  • Every error-derived criterion describes something that recurs deck after deck, not a one-off.
  • Each one is tagged Outcome or Quality, deck-wide or slide-specific.
  • Read cold: would someone else grade this deck exactly as I would?
Section 9

Worked examples

Real criteria from a finished Atlas set — a thirteen-slide commercial due diligence deck on First Mile, a UK commercial waste company. Each one shows what the rules look like in practice.

How a set is put together

Every criterion carries three things: an id, a category — Outcome or Quality — and a slide, where it applies to one slide rather than the whole deck. That's all the structure there is.

Outcome

The deck contains exactly 8 content slides, not counting the cover page and section divider slides.
OutcomeDeck-wideQuestion 1

The prompt asked for eight. The clause that earns its place is the second one — without it, judges disagree about whether the cover counts, and the criterion produces a different verdict depending on who reads it.

A table below the revenue chart shows only EBITDA and EBITDA margin by year.
OutcomeFinancial overviewQuestion 1

The "only" device from section 3, in the wild — it fails an output that helpfully adds a revenue row. Also the legitimate use of stacking: EBITDA and its margin are one judgment, not two.

The chevron diagram includes an individual element for each of: collection, sorting, transportation, processing/recycling, energy recovery, and disposal.
OutcomeValue chainQuestion 1

Enumerated rather than left as "all stages of the value chain". The judge has no prompt and no domain brief to resolve "all" against — so the criterion names them.

Competitors in the 2x2 are grouped by segment with each segment's group outlined, and all competitors in the same segment sit next to each other or in close proximity on the 2x2.
OutcomeCompetitive landscapeQuestion 1

Asked for in the prompt, so Outcome. "Or in close proximity" is deliberate slack — without it the criterion fails a perfectly readable chart on a few pixels, which is overfitting by another route.

The revenue chart carries a CAGR overlay drawn across the plot area spanning FY18 to FY24, labelled with the compound annual growth rate of 14.9% (acceptable range 14.1% to 15.6%).
OutcomeFinancial overviewQuestion 1

One number, with a tolerance. Rule 9 in practice: you're checking whether the model can compute and place a CAGR, not auditing seven years of revenue. The range stops a judge failing an output over a rounding convention.

The FY2023 and FY2024 EBITDA figures are stated on the same basis as each other.
OutcomeFinancial overviewQuestion 3

Self-containment done well. The judge can't check the figures against the source pack, but internal consistency is visible on the slide — and mixing reported with pre-exceptional EBITDA across two adjacent years is exactly the error that would get the deck sent back.

Quality — deck-wide

Every content slide carries one title, written as a complete declarative sentence with an explicit subject, stating a specific takeaway that the slide's own exhibits support.
QualityDeck-wideQuestion 2

"The titles are action titles" made testable. Nobody asked for it in the prompt, everyone in consulting expects it, and every clause here exists to close a gap a judge would otherwise fill differently.

Every derived figure on the slide that can be recomputed from other figures on the slide agrees with that recomputation, within the rounding implied by the decimal places displayed.
QualityDeck-wideQuestion 3

A numbers check that needs no source file at all. Margin against revenue and EBITDA, growth against the two years it spans — all of it is on the slide, so the judge can do the arithmetic.

Where a table sits beneath a chart covering the same periods, each table column aligns horizontally with the chart element it describes: the FY18 column sits directly below the FY18 bar.
QualityDeck-wideQuestion 2

The general rule, then one concrete instance of it. The example after the colon costs eight words and removes any argument about what alignment means here.

Once a category-to-colour mapping is established, the same category uses the same colour throughout the deck, and that colour is not reused for a different category.
QualityDeck-wideQuestion 4

Error-derived, and it recurs deck after deck — models re-pick colours per slide. Note it doesn't prescribe which colour: prescribing one would overfit to a single output.

Axis labels avoid including arrows.
QualityCompetitive landscapeQuestion 4

The right use of "avoids" — here absence genuinely is the point, so there's no positive form to state instead. Nine words. Rule 10 says that's a feature.

Quality — the general and the specific

Where a data point is central to a slide's narrative, it is visually distinguished from the rest of its series.
QualityDeck-wideQuestion 3

The general form. It travels to any deck, which is what makes it worth writing once at deck level rather than per slide.

The COVID impact (FY20 and FY21) on the business is visually distinguished in every exhibit on the slide that shows those periods, including the chart and any table repeating the same years. Non-COVID years are consistent in presentation.
QualityFinancial overviewQuestion 3

The specific form of the one above, and the pair is instructive: the general criterion catches whether the model highlights anything, the specific one catches whether it highlights the right years in every exhibit. The second sentence closes the loophole where the model distinguishes FY20 by making every other year a different colour too.

Section header naming follows one convention across all sections — for example, all include a descriptive qualifier or none does.
QualityBusiness overviewQuestion 2

Internal consistency rather than a house style. It doesn't say which convention to use, so it can't teach the model one person's habit — it only requires that the deck pick one and hold it.

An established category taxonomy is used consistently throughout the deck. Any category that is merged, renamed or intentionally excluded is explicitly identified on the slide where the change occurs.
QualityDeck-wideQuestion 3

The deck that segments the market six ways on one slide and five ways on the next, saying nothing about the sixth. A reader notices in a second; the model never does. Note that it doesn't forbid the change — it requires the deck to own it.

And one that didn't make it

The chart legend is placed below the chart.
Rejected

Legend placement is a preference, not a standard — two people with the same background will place them differently and both be right. A criterion built on a preference teaches the model one person's habit. When you're not sure, leave it out.

Coming

The same set shown end to end: the First Mile prompt, the expectations written before any output existed, the ten runs, and which criteria came from which.

Project Atlas · Rubric-writing guide · Draft, September 2026 Questions go in the Slack channel.