Lab Intermediate Judgment

The Forecast Desk

AI forecasting tools now beat human planners on raw accuracy, so the job shifts from “produce the number” to “frame the question well and decide whether to trust the answer.” You carry one SKU — AuroraGlow serum — from brief to order, switching between briefing the AI and auditing it, and watch a small framing error amplify into the bullwhip you caused. A desk meter rewards a good brief and a real audit — and it can fall.

6 Steps Brief, then audit
~25 min Duration
Trust the answer, on purpose Interactive simulation

Step 1 of 6 · Read the AI’s confident answer

It is the third Monday of the month. The forecast just landed.

You are the demand planner for a skincare brand, mid-way through the monthly planning cycle. One SKU is on your desk: AuroraGlow serum. The AI demand-forecasting tool has already run, and it is confident. Read what it says before you do anything with it.

AuroraGlow serumSKU AGS-30 · the artifact you carry
Recent weekly demand~100 units/wk (steady)
CycleMay plan, 4-week horizon
ForecastIQAI demand engine · auto-run 06:00
“AuroraGlow serum demand is projected to rise +34% next month versus the trailing four weeks. Recommend increasing the replenishment order accordingly. This is a high-confidence forecast.”
+34%Projected uplift
92%Stated confidence
1-click“Approve order”

The number is fluent, decisive, and ships with a green “Approve order” button. That is exactly the trap. Modern AI forecasting genuinely beats human planners on raw accuracy — so your value is no longer producing the number. It is two moves the machine cannot do for you, and you will run both on AuroraGlow across this lab:

1

Brief the machine. Before you ask, frame the inputs and assumptions — horizon, granularity, what is baseline vs. promotion, which external signals matter, and what you are assuming. A vague brief inherits your silos and stale history.

2

Audit the answer. Before you let it drive an order, interrogate the number for accuracy, bias, and business consequence — instead of rubber-stamping it (automation bias) or throwing it out for gut feel (lone-wolf override). The expensive 2026 mistakes are no longer bad spreadsheets; they are these two failures, amplified up the chain as the bullwhip effect.

You will switch between those two stances — briefing and auditing — the whole way through. A desk meter tracks how well you are doing it, and (because this is a real planning cycle) it can fall as well as rise. Do not touch the Approve button yet. First, learn to build the brief the model should have been given.

Stance: briefing the machine

Brief the machine: assemble the AuroraGlow brief

Quality is decided when you frame the question, not when the number comes back. Toggle the elements into your request to ForecastIQ. The brief-quality meter rewards naming the horizon, the granularity, the external signals, and your stated assumptions — and it drops when you leave the brief vague, because the forecast will inherit your silos. This is your desk meter for the rest of the lab; it carries forward.

Brief quality · the desk meter 30%

A bare one-line brief starts you at 30%. Build it up.

This brief’s strength 30%

Two briefs for the same forecast — which is stronger?

Here are two real briefs a planner might send for AuroraGlow. Both ask for next month’s number. Only one will get a forecast you can defend in exec review. Pick the stronger brief, then check your reasoning.

Stance: auditing the inputs

Quality is decided at the input stage, not the output

Your brief framed the question well — now check what that brief actually lets through. A forecast can only be as good as its inputs: garbage in, confident garbage out. Sort each candidate input for AuroraGlow into one of three bins. Click an item to move it: Unsorted → Strong signal → Stale / risky → Out-of-scope → Unsorted.

Strong signal relevant, current, in-scope
Stale or risky real but degraded / siloed
Out-of-scope noise for this SKU/horizon
Candidate inputs ForecastIQ could use
Stance: auditing the answer

You briefed it well. Now audit what it returns.

A good brief does not earn the forecast a free pass. Before a number drives an AuroraGlow order, run it through three lenses, in order. Each one catches a different failure the one before it cannot see.

1

Accuracy

How far off has this forecast been? Measure error against actuals with MAPE or WAPE.

“On average, how wrong?”
2

Bias

Are the misses random, or do they lean one way every cycle — chronic over- or under-forecasting?

“Wrong in which direction?”
3

Consequence

What does that error cost the business — in inventory, fill rate, and cash?

“So what, in units and dollars?”
!

Accuracy alone lies. A forecast can post a respectable MAPE and still be biased — wrong the same way every cycle. MAPE answers “how wrong;” bias answers “which way;” consequence answers “so what.” Skip lens 2 and you will over-order (or under-order) the same direction every month and never see it in the accuracy score.

Lens 1 — how wrong has this forecast been?

ForecastIQ has been forecasting AuroraGlow for a while. Before you trust this month’s +34%, look at last quarter: how far did its weekly forecasts land from the actuals, on average? Slide your estimate of the MAPE (mean absolute percentage error), then reveal it.

Last quarter, ForecastIQ’s weekly AuroraGlow forecasts missed the actuals by an average of how many percent? Lower is better; 0% would be perfect.
Your estimate: 25% MAPE
0% (perfect)25%50% (poor)

Lens 2 — MAPE said good. What is the chart hiding?

Below: ForecastIQ’s weekly forecast for AuroraGlow (gold line) vs. what actually sold (navy line) over the last eight weeks. The 12% MAPE is real. But click every week where the forecast came in higher than actual demand — then check what the pattern tells you.

ForecastIQ forecast Actual demand Click the actual-demand dots that fell below forecast.
Stance: the judgment call

A signal the model never saw just landed

Desk meter · carried from your brief 30%

Your decision here moves it. Two of the three paths move it DOWN.

ForecastIQ’s number for AuroraGlow is +34% at 92% confidence. You now know it runs a chronic over-forecast bias. And this morning a buyer tells you something the model never ingested: your main competitor just went out of stock on their rival serum — their customers are looking for a substitute right now. The model could not have known. What do you do?

Pass the order upstream — and watch the bullwhip

True AuroraGlow demand is dead steady at 100 units/week. But each tier in the chain adds its own safety buffer on top of the order it receives — so a small error at the retail shelf gets amplified at every hop upstream. Pick which forecast the retailer starts from, then run the three rounds and watch the order swing grow. This is the bullwhip effect, and in the first mode it is the one you caused — running it to the end drops the desk meter.

Retailer

orders from distributor
100
true demand: 100

Distributor

+ buffer, orders from factory
100
should be: 100

Factory

+ buffer, builds capacity
100
should be: 100

Brief well, audit always

1. Where does AI sit in the S&OP cycle?

The monthly Sales & Operations Planning (S&OP) cycle is shuffled below. Put it in order. The two placements that matter most: “AI generates the baseline forecast” belongs inside the cycle, not above it — and “planner audits & overrides” comes right after it, before the number is ever committed.

    2. Justify the final number to exec review

    This is where the cycle ends: the order goes up for sign-off and someone asks “why this number?” Write the two-to-three-sentence note you would attach — whether you kept or adjusted the forecast — citing accuracy, bias, and the external signal. Then self-assess against the checklist and reveal a model note.

    Self-assess: does your note…

    3. Field guide

    One page to keep. Everything the Forecast Desk taught, at a glance.

    1Brief, don’t just ask

    • Pin the horizon + granularity
    • Separate baseline from promo lift
    • Name the external signals
    • State assumptions + a re-run trigger

    2Input quality decides everything

    • Recent, SKU-specific, in-scope = strong
    • Stale history & stockout zeros = risky
    • A siloed plan the model never saw = garbage-in
    • “AI can” ≠ “AI should use this”

    3The three-lens audit

    • Accuracy: MAPE/WAPE band (<10 excellent, ~20 good, >30 intervene)
    • Bias: are all misses one direction?
    • Consequence: units, fill rate, cash
    • Accuracy alone hides bias

    4Trust, tune, or override

    • Blind trust = automation bias (understock)
    • Gut override = lone-wolf bias (dead stock)
    • Add the signal & re-run = bounded, defensible
    • The skilled move is in the middle

    5Mind the bullwhip

    • Steady demand, amplified swings upstream
    • Your un-audited error becomes the factory’s
    • An audited number flattens the swing

    6AI sits inside S&OP

    • AI generates the baseline
    • The planner audits & overrides before commit
    • Never put AI above the process

    You ran the forecast desk.

    You briefed the machine so it couldn’t inherit your silos, audited the answer for the bias its accuracy hid, made the trust/tune/override call with consequences, and watched a small framing error become a bullwhip when the audit was skipped. That two-stance discipline — brief well, audit always — is the planner’s job now that the machine produces the number.

    Step 1 of 6