The first artefact a delivery program traditionally produces is the status report. Milestones. Schedule variance. Effort burn. Risk log. A green / amber / red column down the right side that lets a steering committee scan the page in thirty seconds.
I had built that page for two decades. I knew which boxes to bold, which numbers to lead with, which risks to surface and which to keep on the watch list. The status page was the unit of leadership conversation.
In the retail transformation program, I noticed the status page was getting fewer and fewer questions.
The questions that replaced it
What leadership was actually asking, by the end of the first month, sounded different.
"How confident are we?" — meaning, what is the probability we hit the holiday deadline at the quality we promised?
"Why did the AI change the approach?" — meaning, the coding agent re-ordered the queue last week, walk me through the logic.
"What is the cost trend?" — meaning, consumption is volatile, what does the curve look like and where is it going?
"Can we trust the outputs?" — meaning, the testing agent passed three thousand cases this week, do we believe the green?
"What are my options?" — meaning, here’s a decision I might need to make, lay out the alternatives and the trade-offs.
None of these questions are answered by milestones, schedule variance, or a red / amber / green column.
The artefact that does answer them
The reporting artefact I now build for AI-bearing programs has four sections, not the traditional seven.
Confidence band. Not "we are 78% complete." Something more like "we are 80% confident we hit the holiday deadline at the quality bar; 95% confident we ship something working; 50% confident we hit the optimistic cost target." Confidence comes from the program’s own evaluation harness plus human judgment from the leadership team. It is updated weekly.
Workflow health. A small dashboard, not a paragraph. For each major workflow: agent eval scores, human gate throughput, drift signals, consumption run-rate. Anything red gets a sentence underneath. Anything green stays a number. Read it as one paired signal, agent eval scores tell leadership whether the agents are driving well; human gate throughput tells them whether the human navigators are keeping up. Either one in the red is a problem. Both in the green is what running well looks like.
Decisions on the table. The page that earns its keep. Each decision laid out as: the question, the options, the trade-offs, the recommendation, the deadline. Leadership comes to the steering committee to make decisions, not to receive status. This is where they make them.
Outcome trajectory. Where the program is heading, not where it has been. Are we trending toward the original outcome or away from it? Trending faster or slower? With a single chart, not a paragraph.
How this lands in the room
The first time I took a confidence-band report to a steering committee, the CFO asked me where the percent-complete number was. I said the program no longer had one — we had a confidence band, and here was what fed it. Three minutes of explanation. He asked two follow-up questions. By the third committee meeting, the confidence band was the metric the room asked for first.
The mistake to avoid is leaving the old artefacts in place "for compatibility" while introducing the new ones. Two reports in parallel teaches leadership that the new one is optional. One report, new shape, with a footnote pointing to the old numbers if anyone needs them — that’s the path that takes.
Why this matters more than it sounds
The status page wasn’t neutral. It taught leadership to ask retrospective questions ("are we on track against last month’s plan?") instead of prospective ones ("what should we decide this month?"). In a deterministic delivery, that was fine. In a probabilistic delivery, it was actively harmful — leadership was being trained to look in the wrong direction.
Reporting is a behaviour shaper. Changing the artefact changes what gets discussed. Changing what gets discussed changes what gets decided. Most of the lift in AI-first leadership reporting is downstream of the artefact swap.
Status decks teach retrospective questions. Confidence reports teach prospective ones.
A note on dashboards
The four-section reporting artefact above is the steering committee version. The live operational version sits in a dashboard — agent evals, gate throughput, consumption, drift — that the program team watches continuously. The steering committee gets a curated read of the dashboard, plus the decisions and the trajectory. Both layers exist. Both are needed.
If you only build the dashboard, leadership tunes out. If you only build the deck, the program runs blind between Wednesdays. The pair is what works.
The starting move
If you run a steering committee next week, rewrite the cover page. Lose the percent-complete. Replace it with a confidence band, with three lines underneath explaining what fed it. Watch what gets asked. The room will tell you very quickly whether the new artefact has caught.
In my experience the new artefact catches inside two committees. By the third one, you are asked for it.