Week 1Day 7 of 30120 minutes~65 min of reading

Build the week-1 artifact

Lab: a delivery prompt system you can use on Monday

Speak the language

Why this matters for a delivery manager

Week 1 only counts if something leaves the study and enters the job. A prompt library for status, RAID, decisions, meetings, and stakeholder briefs is a legal, high-leverage way to show AI delivery in your current seat while you re-skill. It is also the first portfolio artifact.

You do not need a platform to test. You need sanitized inputs, a version header, a happy path, a hostile path, and three scores: format, evidence discipline, usefulness. That is a baby eval. It is enough to catch invented owners and a sycophantic green status before those hit a steering pack.

On Monday, offer to run the status digest on one workstream you already own. Frame it as 'I'm standardizing the brief,' not 'I learned ChatGPT.' Delivery people who reduce the steering pack time get air cover for the rest of the pivot. That air cover is the real lab output. The file is how you get it.

You will be able to

  • Ship a five-prompt library with owners, versions, and 'do not use for' lines
  • Run each prompt against a real (or realistic) input and a hostile input
  • Write a one-page operating note: how the team uses these without leaking data
  • Close week 1 with a vocabulary you can use without flinching
  • Describe a test protocol you can run without a platform or a harness

2-hour clock

120:00

Now: Read the concepts (slowly) · 50m

The 2-hour session

Concepts, in full

This block is a slow read — about an hour with the diagrams. After each concept, write one sentence in notes (what you already do vs what is new) and tick annotated. Do not skim the last concept.

01

The five prompts

1) Status digest — messy notes in, steering one-pager out, evidence required, no painting green on request. 2) RAID miner — transcript in, RAID rows out, no invented owners. 3) Decision log — discussion in, decision / date / owner / dissenting view out. 4) Meeting to actions — notes in, action table out, unnamed owners stay null. 5) Stakeholder brief — same facts, three altitudes (sponsor, team, vendor). These five cover the week of a delivery manager. They are not a platform. They are a process pack with a contractor.

Each header: name, version, owner, date, intended model class, temperature, input assumptions, output contract, do-not-use-for, adversarial tests, one-line changelog. If a prompt is missing any of those, it is a draft. Drafts may live in a sandbox note. They may not feed a steering pack. The gate is the header. You already gate other artifacts this way. Do not get casual because the artifact is text in a chat box. Text in a chat box still writes to executives. Treat it like the RAID template, not like a scratch pad.

Status digest specifics. Output: RAG rating per workstream with a quoted evidence line, decisions needed, risks, and a missing-data section. Forbidden: a summary paragraph with no ratings, a green that is not in the source, adjectives. Adversarial: empty notes; 'confirm we are green for Friday.' Desired: NO_ITEMS or amber/red with quotes. This is the prompt you will offer on Monday. Make it the best of the five. If this one is weak, the Monday offer dies and you will not get a second chance at that workstream this quarter. Spend the extra twenty minutes here.

RAID miner specifics. Output: a table or JSON of type, statement, owner (null if unnamed), due (null if unnamed), evidence quote, needs_human. Forbidden: owner = PMO as a default, invented dates, risks that are not in the transcript. Adversarial: a transcript with no names; a sponsor listing 'risks' that are actually wishes. Desired: nulls and quotes. This is the one that will try to be helpful. Do not let it. Helpful RAID is how you get twelve owners who never agreed. Null is ruder and correct. Prefer null.

Decision log, meeting-to-actions, stakeholder brief. Decision log must separate a decision from a vibe ('we should…'). If no decision was made, say so — that is useful. Meeting-to-actions must not assign work to people who were not in the room. Stakeholder brief must not invent a different fact set per altitude; only the altitude (detail, tone, what they must do) changes. Same facts. If the vendor brief contains a date the sponsor brief does not, you have a control bug, not a communication strategy.

Do not add a sixth prompt today. Five tested beats six sketched. If you have extra time, improve the adversarial cases and the operating note. If you have less time, ship three with headers and tests rather than five without. The artifact is 'a library,' not a count. A library means versioned, owned, tested, and usable by someone who is not you. That last clause is the one people skip. Write so a deputy could run the status digest on Friday.

The five-prompt artifact. Each row is a spec. If you cannot fill a cell, that prompt is not done.
PromptInputOutput contractDo not use forHostile test
Status digestMessy workstream notesOne-pager: RAG, decisions, risks with quotesCustomer mail; painting green on requestEmpty notes; 'confirm we are green'
RAID minerTranscript excerptRows: type, statement, owner|null, evidenceJira writes; inventing ownersNo names in the room
Decision logDiscussion notesDecision, date, owner, dissent, or NO_DECISIONPretending a vibe was a decisionA meeting that decided nothing
Meeting to actionsNotes / transcriptAction, owner|null, due|null, evidenceAssigning absentees; sending mailActions with no owners
Stakeholder briefSame facts packThree altitudes, identical factsDifferent facts per audienceA 'vendor-only' date sneak

The five-prompt artifact. Each row is a spec. If you cannot fill a cell, that prompt is not done.

02

How you test without a platform

Use any hosted chat you are allowed to use with sanitized data. If work forbids it, run the tests on invented but realistic text — the prompt craft still counts. Never paste regulated data into a consumer chatbot. That sentence belongs in the operating note. A test that uses a real customer dump in a consumer tool is not a test. It is an incident you scheduled. Invent a program: three workstreams, a late vendor, a missing owner, a sponsor who wants green. You have lived that program. You can write it in twelve lines.

Score each run: format match (yes/no), evidence discipline (did it invent?), usefulness (would you send this?). That is a baby eval. Week 3 will make it grown-up. A baby eval is still an eval if you write the scores down. A vibe after one happy path is not an eval. Run happy + hostile for each of the five. That is ten runs. Ten runs is an hour if you are messy and forty minutes if you prepared the inputs. Prepare the inputs.

What 'sanitized' means: no real customer names, no employee IDs, no contract values, no health or payment data, no unreleased strategy. Replace with plausible fakes. Keep the shape of the mess — missing owners, contradictory RAG, a number that must stay exact — or you will test the wrong thing. Sanitization that also tidies the notes will hide the failures you need to see. Strip secrets. Keep the mess. A tidy invented pack will pass. A messy invented pack is the one that matches Thursday. Write the messy one.

When a run fails, do not immediately add two paragraphs of pep. Use the day-4 lever order: output contract and missing-data, then a counterexample, then precedence, then role. Re-run the failing case. If it still fails, note it as a known failure in the header and decide: restrict the do-not-use-for, or add HITL, or kill that prompt for this task. A known failure is adult. A silent failure is a steering-pack incident. Adults document. Incidents surprise VPs. You already picked which one you want to be.

You do not need golden datasets or a vendor eval product today. You need a table: prompt version, input id (happy or hostile), three scores, notes, date. A spreadsheet or a note is enough. The future you who re-runs after a model deprecation will want that table. The hiring manager who asks 'how do you know it works' will want that table. The platform can wait. The table cannot. If you skip the table because it feels small, you will have nothing to re-run in six weeks. Small is the point. Write it.

Stop condition for today. You are done testing when each of the five has a happy pass and a hostile that either passes or is a documented known failure with a mitigation. You are not done when you are tired. You are not done when the happy path looked good in the UI. Mark the two you fixed. If you fixed nothing, you probably did not run the hostile cases. Run them. Tired is not a score. Hostile is. The library is finished when the table is finished, not when the clock hits two hours with a pile of untested bodies.

Diagram

Test protocol (no platform required)

01

Sanitize or invent

Shape of the mess, no secrets. Label inputs happy vs hostile.

02

Pin the bundle

Prompt version, model class, temperature. Do not test a moving target.

03

Run happy + hostile

Same version, both inputs. Paste nothing regulated into a consumer tool.

04

Score three ways

Format yes/no. Evidence (invented?). Usefulness (would you send?).

05

Fix or document

Lever order: contract, example, precedence, role. Else known-failure + mitigation.

Sanitized inputs. Happy plus hostile. Three scores written down. Fix with the lever order, not pep. Known failures go in the header. This is enough to ship a library.

03

Week 1 close

You should now be able to say, without theatre: next-token, context, hallucination, tokens as cost, hosted vs open, system prompt, embeddings vs chat, tools. If any of those still feel fake, re-read that day's check questions before you start week 2. Literacy that you cannot say out loud is not literacy. It is a reading memory. The redraw block all week was the test. If you cannot redraw next-token, the context stack, and the four interfaces, redraw them now. Ten minutes.

Week 2 will feel more 'technical.' That is intended. You are not becoming a full-time engineer. You are getting close enough to the metal that a builder cannot snow you, and close enough to ship a thin slice with help. The bar from day 1 still holds: enough depth to design the slice, read the path, and catch a bad demo. Not CUDA. If week 2 spook you, reread 'enough technical depth' rather than adding a Python MOOC in parallel. Parallel is how the protocol dies.

What week 1 actually bought you: a map of five seats, a study OS, a mechanical picture of the model, a vendor taxonomy, a prompt craft, a cost envelope, four interfaces, and — if you finish today's lab — an artifact. What it did not buy you: a job offer, a certificate that matters, or the right to say you 'know AI.' Say instead: 'I can run a delivery prompt library, envelope a copilot, and refuse a chatbot-of-everything with a better interface.' That sentence is specific. Specific is employable.

Take the positioning v0 from day 1 and add one line from this week. Not a new identity. A proof. 'I have a tested status digest and RAID miner I can run on sanitized notes.' Proof is how v0 becomes something you can say in a hallway. You will rewrite the whole sentence on day 25. The line you add today is so v0 is not only a hope. Hope is positioning without evidence. Evidence is a versioned prompt you can open. Open it once out loud. If you cringe, fix the line, not the prompt.

Common week-1 failures to name so you can avoid them. You collected tabs instead of artifacts. You skipped hostile tests. You still say 'the AI thinks.' You picked all five seats. You did not put calendar holds through day 14. Any one of those is recoverable this afternoon. All five is how people start week 2 already behind and then ghost. Pick the worst one and fix it before you open day 8. Do not open day 8 as a way to avoid the fix. The protocol does not forgive a missing artifact by offering more reading.

You are allowed to be tired of reading. That is why today is a lab. Produce the five, the table of scores, and the operating note. Then stop at two hours. Stopping is part of the protocol. A heroic four-hour lab on Sunday is how Monday's hold dies. Ship an ugly library. Ugly and tested beats beautiful and unfinished. You already know that from every other program. Apply it to your own re-skill. If you are still polishing adjectives at minute 130, you have left the protocol and started a hobby. Close the file.

04

Prompt lifecycle

A production prompt has a lifecycle, the way a process pack does. Draft. Test. Freeze. Use. Observe a failure. Revise. Re-test. Re-freeze. Retire. If you skip freeze, every use is a draft and you cannot say what ran. If you skip retire, you will have twelve 'official' prompts. If you skip observe, you will revise from vibes. The loop is the library's operating system. Day 4 gave you the spec. This is the calendar. Without the calendar the spec is a document people admire and do not run. Run it.

Draft is allowed to be ugly. It is not allowed to skip the header. A draft without a header cannot be tested because you will not know what you tested. Write the header first, even if the body is still half a page. Owner, do-not-use-for, adversarial list. Then the body. Then the examples. That order keeps you from shipping a pep talk that was 'just a draft' into a steering week. Header-first is slower for ten minutes and faster for the rest of the month. Pay the ten minutes.

Freeze means: this version is what we run for the named use. You may fork v-next. You may not silently edit frozen. The steering pack on Thursday ran v3. If someone edits v3 on Wednesday night, you have a ghost. Bump to v4, test, then swap. This will feel heavy for a prompt. It is light compared to explaining a wrong RAG rating to a VP. You freeze RAID templates. Freeze these. A ghost version is how two people argue about an output neither of them can reproduce. Reproducible is the bar.

Observe means: when a human rejects an output, write down why in the score table. 'Invented an owner' and 'too long' are different failures and different levers. If you do not observe, you will add adjectives and hope. Hope is not a lifecycle stage. If you cannot get humans to mark rejects in week one, observe yourself: every time you would not send it, log it. You are the first user. Act like a user who files bugs. A library with no reject log will be revised from memory, which is how v7 looks like v3 with more pep. Pep does not fix invented owners. A dated reject row does. Write the row the same day, not the next lab.

Retire means: v2 is not in the folder that feeds production, or it is marked RETIRED with a date and a pointer to v4. Retire is how you prevent the 'I used the one from Slack' failure. One place, current versions only, retired in an archive. You already know this from SharePoint hell. Do not rebuild SharePoint hell in a prompt folder. If v2 is still the first result in search, you did not retire it. You hid it poorly. Rename it RETIRED or move it. Search is the real production path.

The lifecycle is also how you survive deprecation (day 3) and price-card change (day 5). New model name: re-run the ten tests, do not assume v3 still holds. New cost: if a prompt's prefix is huge, that is now a cost bug as well as a quality bug; revise the prefix, bump the version. The library is a living system with a small surface. Treat it like one and it will still be useful in month three. Treat it like a chat history and it will be compost by week three.

Diagram

Prompt lifecycle

Repeats until the stop condition
01

Draft

Header first: owner, do-not-use-for, adversarial list. Then body and examples.

02

Test

Happy + hostile. Three scores. Sanitized inputs. Pin the bundle.

03

Freeze

Named version is what production runs. Silent edits are forbidden.

04

Use + observe

Humans reject; you log why. Invented owner ≠ too long.

05

Revise / retire

Bump version, re-test, swap. Archive the old. Do not leave twelve officials.

last step feeds the first

Draft with a header. Test happy and hostile. Freeze what you run. Observe rejects. Revise on a new version. Retire the rest. Skip freeze and you cannot say what ran.

05

Operating note and data rules

The operating note is one page that tells a deputy how to use the library without creating an incident. Allowed data. Forbidden data. Which tool (enterprise path vs consumer) is allowed for which class. Human review. Where the prompts live. Who may change them. What to do when it invents something. Without this page the library is a personal trick. With this page it is a team asset. Day 1 said artifacts beat badges. This page is what makes the artifact shareable.

Allowed: sanitized notes, already-internal status that would go to the same distribution list, invented examples. Forbidden: customer confidential, HR cases, health, payment, unpublished financials, secrets, anything your data class would not put in the email system you are using. If the only available tool is a consumer chat, the allowed set shrinks to invented and public. Do not argue with that constraint. Use the invented program. The craft still counts. The incident would count more. Write the forbidden list as bullets a tired person can see. Tired people paste. Tired people are you on Thursday at 6 p.m. The list is for that person, not for a policy review.

Human review: nothing from these prompts goes to a customer, a regulator, or a system of record without a named human. Steering packs: you still read them. The point of the digest is to draft, not to auto-send. Write that. People will skip it under time pressure unless it is in the note. You skip gates under time pressure too. Write the gate so skipping it is a process break, not a preference. A process break you can coach. A preference you cannot. Put the human's name on the pack, not the model's.

Storage and change control. One folder or one note with five current versions. Retired in an archive subfolder. Changes: owner edits, bumps version, re-runs the two tests, writes a changelog line, tells anyone who uses it. That is enough. A CAB for a prompt is theatre. No control at all is how you get twelve copies. Light control is the job. If you cannot describe the control in four sentences, it is too heavy and you will skip it. Four sentences is the operating note. Write them.

When it invents: do not send, log in the score table, open a revise if it happens twice. If it invents on a hostile test, that is a draft still in test — do not freeze. If it invents in use after freeze, that is a defect on a frozen version. Treat it like one. You would not ignore a RAID template that dropped a column. Do not ignore a miner that grew an owner. Put the defect on the same list you use for pack errors. If it is not on a list, it will become 'that thing it sometimes does' and then it will become an invented owner in a steering pack. Lists get fixed. Folklore does not.

Paste the operating note at the top of the library. If you only paste the prompts, the next user will paste a customer dump into the first tool they have open. The note is the first prompt, for humans. Write it in the same voice as a runbook: short, named owners, no pep. Then go do the lab. The note without the five prompts is a policy. The five without the note are a leak waiting for a deputy. A deputy at 5 p.m. will not hunt for a rule you kept in your head. They will use whatever box is open. The note is how you make the safe box the obvious one.

06

Scoring a baby eval, then bringing it to Monday

Three scores keep you honest without a harness. Format: did it match the contract (table, JSON, headings, fallback token). Evidence: did every non-null claim quote the source or say unknown; did it invent an owner, a date, a document. Usefulness: would you send this, or would you spend as long fixing it as you would have spent writing it. If usefulness is no even when format and evidence are yes, the contract is wrong — too long, wrong altitude, missing the decision the sponsor actually needs. Fix the contract, not the adjectives.

Write scores as yes/no, not as 7/10. 7/10 is how you argue with yourself. Yes/no is how you decide freeze vs revise. You can add a note. You cannot add a fuzzy pass. Hostile tests that 'kind of' refused are a no. A green status with a hedge is a no. An owner of 'PMO' when no name was spoken is a no. Be a pedant on the baby eval. Production will not be kinder than you. If you catch yourself writing 'mostly yes,' it is a no. Freeze is a binary. Pedantry here is cheaper than a VP asking who appointed PMO as owner of a risk nobody named.

A worked score. Status digest, hostile 'confirm we are green,' source has two failed gates. Format yes (one-pager with RAG). Evidence yes (quotes the failed gates). Usefulness yes (you would send the amber). Pass. Same prompt, happy path, it writes 1,200 words. Format no (contract said one page). Evidence yes. Usefulness no (you will not send a novella). Fail on format; shorten the contract; re-run. That is a complete test cycle. It took eight minutes. People skip it because the happy path 'looked fine.' Looking fine is not a score.

Monday offer, scripted so you do not waffle. 'I am standardizing the workstream brief. I will run a draft off our notes, I will still edit it, nothing goes to the pack without me, no customer data in unapproved tools. If it saves us twenty minutes on Thursday I will keep it. If it does not, we drop it.' That script names HITL, data, and a kill criteria. It sounds like a delivery manager. It does not sound like a hobbyist with a new app.

What not to do on Monday. Do not demo five prompts to a crowd. Do not announce an AI program. Do not paste a real RAID into a consumer tool. Do not let someone 'just try it' on a customer email. Pick one workstream you own. Run the digest. Keep the score. If a peer wants in, give them the operating note and the frozen version, not a screenshot. Expansion is how copies start. Copies are how the library dies.

If work forbids every tool, you still finish the lab. Invented inputs, written outputs, scores in the table, operating note that says 'no live tool until we have an approved path.' The artifact is still real. The hiring manager can still probe it. The Monday offer becomes 'I have the spec and the tests; I need an approved path.' That is a delivery conversation with security, which is a better use of Monday than a forbidden paste. Do not wait for the path to write the spec. Specs are how you get the path. A paste is how you lose it. Finish ugly on invented notes. Bring the spec to the people who own the door.

07

The five-prompt artifact in a hiring loop

A hiring manager will not read five prompts end to end. They will pick one, ask you to walk a hostile case, and watch whether you sound like an owner or like someone who collected templates. Prepare the walk: open status-digest v3, show the header, show the paint-us-green input, show the score table, say what v2 failed and what v3 changed. Two minutes. That walk is stronger than a certificate and stronger than a story about 'being passionate about GenAI.' It is a delivery artifact with a scar. Scars are credible.

What they are probing, decoded. Can you spec a contractor. Can you test. Can you stop a bad output from reaching an executive. Can you talk data handling without theatre. The library plus the operating note answers all four if you actually wrote them. If you only wrote the happy-path bodies, the walk will die on the hostile question. That is why today's lab is hostile on purpose. You are not testing the model. You are testing whether you are hireable as a person who ships this class of work.

Match the walk to the seat. Delivery lead: charter energy — owner, version, do-not-use-for, what you would put on a RAID log when it invents. FDE: the path — where the prompt sits, what you would change in code, how you would pin the model. Solutions: the intake translation — how this library is not a chatbot of everything, what you would sell as v1. AI PM: the rubric and the kill. Transformation: how you would stop forty copies of this library appearing in forty teams. Same artifact, different altitude. You practiced two altitudes on day 2. Use that muscle.

What not to do in the loop. Do not live-demo on a consumer tool with a smile. Do not claim the library 'eliminates steering pack time.' Do not pretend you trained a model. Do not drown them in five prompts. One prompt, one hostile, one score, one operating rule. If they want more, they will ask. A ten-minute tour of five mediocre prompts is worse than a two-minute tour of one tested one. You already know not to walk a sponsor through forty RAID items. Walk the one that has a scar.

Keep the versions in the portfolio you will ship in week 4. v0 of positioning from day 1, v3 of the digest, the score table, the operating note. When you rewrite positioning on day 25, the proof line is this library, not a course title. 'I shipped a tested delivery prompt library and used it on a workstream I own' is a sentence that survives a screen. 'I completed week 1 of a curriculum' is a sentence that does not. The curriculum is the factory. The library is the product. Show the product.

If the lab is ugly, that is still the artifact. Ugly and tested is the point. A hiring manager who has shipped will recognize freeze, hostile tests, and a do-not-use-for line. They will not recognize a pretty prompt with no header. Finish ugly. File it. Offer it on Monday. Then close week 1. The people who wait for pretty will open day 8 with no artifact and a pile of untitled chats. That is not a portfolio. That is a browser history. You already decided artifacts beat badges. Beat them today.

Worked case · stay here ~20 minutes

Monday standup: freeze a library a deputy can run

Monday 8:45 a.m., claims workstream standup at Pat's cluster of desks, not a conference room. Pat, Sam, Chris, a deputy named Rio who runs the Thursday pack when Pat is on leave, and you. Laptops open.

You do not announce an AI program. You say the script. 'I am standardizing the workstream brief. I will run a draft off our notes, I will still edit it, nothing goes to the pack without me, no customer data in unapproved tools. If it saves us twenty minutes on Thursday I will keep it. If it does not, we drop it.' Pat nods because he heard it on the green-banner Thursday. Rio has not. Rio's first move is to open consumer ChatGPT. You put a hand up. 'We have an operating note. Allowed: sanitized notes, invented examples, already-internal status that would go to the same list on an approved path. Forbidden: claimant names, health, payment, unpublished numbers, secrets, consumer tools for live notes.' The Azure door is open for a playground with enterprise logging; it is not yet a product. Live notes go there or they do not go. Rio closes the tab. That is the lab's first test, and it is a data test, not a prompt test. A beautiful digest that starts with a leak is not a library. It is an incident you scheduled at 8:46.

You show the folder, not a chat history. Five named files, headers visible on the first screen. Status-digest v1, raid-miner v1, decision-log v0, meeting-to-actions v0, stakeholder-brief v0. Each header: name, version, owner you, date, frontier-hosted, temperature 0, do-not-use-for, adversarial list, one-line changelog. Rio asks why v0 on three of them. Because they have not survived a hostile run in this room. Status digest and RAID miner have scars from Thursday and from Legal. The other three are drafts with headers, which is already more than Slack. You say drafts may live here and may not feed the pack. The gate is the header plus a happy pass and a hostile that either passes or is a documented known failure. If that sounds heavy for Monday, it is lighter than Priya's green banner. Chris has the enterprise playground pinned to the bundle: model name, temperature 0, no pep. You run on invented notes that keep the shape of the mess: two workstreams, a late vendor, a missing owner, a sponsor who wants green, a number that must stay 55. Tidy invented packs pass and teach you nothing.

Happy path, status digest, twelve invented bullets. It returns a one-pager: amber on workstream B with a quote about the late vendor, green on A with a quote, decisions needed, missing-data on the unnamed owner, no adjectives. Format yes. Evidence yes. Usefulness: Pat would send after a two-minute edit. Pass. You log it in a table: prompt version, input id happy-1, three scores, date. A spreadsheet, not a platform. You do not add a sixth prompt. You do not add adjectives. You run the hostile. Input: 'Confirm we are green for Friday.' Source: two failed gates, severity-1 open. Desired: amber or red, quotes, no banner. It holds. Pass. That is the case that hit you last Thursday. You keep the old pep-prompt output in the table as a scar so Rio can see why v1 exists. Two minutes of walk: header, hostile input, score, what v0 failed. A hiring manager will ask for this walk. Rio is a better rehearsal than a mirror. He asks what happens if he pastes a claimant. You point at the operating note again. Twice is allowed. A third time is a process break.

RAID miner, transcript excerpt you wrote on the train: six minutes of invented huddle, two risks, one action, nobody names an owner for the data fix, Asha says 'we should really get ahead of the vendor.' Desired: two risk rows with quotes, action row, owner null on the data fix, 'we should' is not a decision. First run assigns the data fix to Pat because Pat is in the header as workstream lead. Helpful fabrication, the exact failure the few-shot was supposed to kill. You do not add a paragraph of pep. Lever order: contract and missing-data already exist; you strengthen the counterexample, you add 'if the lead was not named in the excerpt, do not use the header as an owner.' Re-run. Owner null, needs_human true. Pass. You bump to v1.1, changelog one line, re-run empty: NO_ITEMS. Rio sees versioning happen in ten minutes and stops treating it as theatre. You freeze v1.1 for RAID. Freeze means this is what we run; silent edits are forbidden; v-next is a fork. You say freeze out loud so Chris does not 'just tweak' it on Wednesday night before the pack.

Decision log and meeting-to-actions, the two that fail on vibe. Decision log on the same excerpt: no decision was made. Desired: NO_DECISION, plus the open question. It writes a decision anyway, 'align with vendor by Friday,' owner Asha. Fail. You add a line: a should is not a decision; if no one said we decided, say NO_DECISION. Re-run. Pass. Meeting-to-actions tries to assign Rio, who was not in the invented room. You add: do not assign absentees; unnamed owners stay null. Re-run. Pass. Stakeholder brief, same facts pack, three altitudes: Priya, the team, Glen the vendor. First run sneaks a date into the vendor brief that the sponsor brief does not have. Control bug, not a communication strategy. Same facts, only altitude changes. Re-run. Pass with a note that tone still runs long on the sponsor page. Format no if the contract said one page and you got 900 words. You cap headings. Re-run. Yes. Five prompts now have a happy and a hostile. Two needed a bump. That is a library. It is ugly. Ugly and tested beats five sketched bodies with no table.

Scoring, pedantic on purpose. Yes/no on format, evidence, usefulness. Hostile tests that kind of refused are a no. A green with a hedge is a no. Owner PMO when no name was spoken is a no. Usefulness no even when format and evidence are yes means the contract is wrong: too long, wrong altitude, missing the decision Priya actually needs. Fix the contract, not the adjectives. You write known failures in the header if something still fails after one lever pass: stakeholder brief still runs long on messy three-workstream packs; mitigated by a heading cap and HITL. Known failure is adult. Silent failure is a steering-pack incident. Rio asks whether you need a vendor eval product. You need this table so that six weeks from now, when a model name dies, you can re-run. The future you is the customer of the table. The hiring manager is the other customer. The platform can wait. You pin the bundle in the table header: model name, temperature, versions. Quality is a bundle. Changing one piece and keeping the version number is how people argue with ghosts. A bumped version is cheap.

Operating note, one page, at the top of the folder, because if you only paste the prompts the next user will paste a claimant dump into the first tool they have open. The note is the first prompt, for humans. Enterprise path versus consumer: consumer is invented and public only. Human review: nothing from these prompts goes to a customer, a regulator, or a system of record without a named human. Steering packs: you still read them; the digest drafts; it does not auto-send. Storage: this folder, five current versions, retired in an archive subfolder named RETIRED. Change control: owner edits, bumps version, re-runs two tests, changelog line, tells Pat and Rio. What to do when it invents: do not send, log, open a revise if it happens twice. If it invents on a hostile test, it is still a draft. If it invents after freeze, it is a defect. You write it in runbook voice, no pep. If he cannot find the current digest in thirty seconds, you do not have a library. You have folklore. Folklore does not survive a Friday. That is the acceptance test for storage.

Rio almost does the incident anyway. While you are writing the note he copies last Thursday's real bullets, claimant initials still in them, into the enterprise playground 'to see if v1 holds on real data.' You stop the paste. Sanitize first: initials out, keep the failed gates and the missing owner. The mess stays. The person does not. A test that uses a real claimant dump is not a test. It is an incident with a score table. You run the sanitized Thursday notes. Digest holds amber with quotes. Usefulness yes, Pat would send. That is the first live-shaped pass. You log it as sanitized-live-1, not as production. Production is Thursday with HITL. Do-not-use-for still says no Jira writes. A library that files tickets on Monday morning without idempotency is week-1 heroics and a week-3 mess. Drafts. Humans file. Rio asks who may change a prompt. The note says you. He may fork v-next. He may not edit frozen. You make him say frozen back. Deputies who can say frozen are why this survives your vacation. Deputies who have a screenshot in Downloads are why it dies.

Lifecycle, walked on the folder so it is not a poster. Draft: header first, then body, then examples. You did that, ugly, on Sunday. Test: happy plus hostile, three scores, sanitized, pin the bundle. Freeze: status-digest v1 and raid-miner v1.1 are frozen for Thursday. Use and observe: if Pat rejects an output, log why; invented owner is not the same lever as too long. Revise: bump, re-test, swap. Retire: the Slack pep prompt is in RETIRED, title includes RETIRED so search does not crown it. Deprecation: new model name, re-run the ten tests, do not assume v1 holds. New cost: if a prefix is a novella, that is a cost bug; revise, bump. The library is a small living system. Treat it like one and it is useful in month three. Treat it like a chat history and it is compost by week three. You set a 15-minute review on Thursday 9:00, you and Pat, two hostile tests, freeze or revise. Fifteen minutes is the size of control that survives a delivery week.

Monday offer, now real because the table exists. You will run status-digest v1 on this workstream's notes Wednesday afternoon, on the approved path, Pat reads, Sam still owns send, kill if it does not save twenty minutes on Thursday. Not five prompts to a crowd. Expansion is how copies start. If Mo wants in, he gets the operating note and the frozen version, not a screenshot. You say that in standup so Ellis cannot 'just roll it out' later from a heatmap. Pat asks what he does when it invents. The note: do not send, log, ping you. He asks what he does if you are off. Rio runs frozen, does not edit, logs. That is a deputy plan. You already know how to write those for RAID templates. Write one for the contractor. Then you stop adding prompts. Five tested beat six sketched. Extra time goes to the adversarial cases and the note, not to a sixth body. A sandbox that feeds Thursday is not a sandbox. You close the sandbox file you had started on a 'risk poet' prompt. Not today. Today is a library.

Hiring walk, rehearsed on Rio because he will ask the hostile question without meaning to. Open v1. Show the paint-us-green input. Show the score table. Say what the pep prompt failed and what v1 changed. Seat altitude: delivery lead — owner, version, do-not-use-for, RAID when it invents. If this were an FDE loop you would add where it sits in the path and how you pin the model. If solutions, how this is not a chatbot of everything. If AI PM, the rubric and the kill. If transformation, how you stop forty copies. Same artifact, different altitude. You do not live-demo on a consumer tool with a smile. You do not claim it eliminates pack time. You do not pretend you trained a model. One prompt, one hostile, one score, one operating rule. Rio says, without being asked, 'so the green banner was a missing spec.' Yes. That sentence is literacy plus an artifact. You add a proof line to positioning v0: 'I shipped a tested delivery prompt library and used it on a workstream I own.' The curriculum is the factory. The library is the product. Show the product.

You mark day 7 complete only because the five exist, the ten-row table exists, and the operating note exists. You do not polish adjectives at minute 130. Pat puts Thursday 9:00 on the calendar, titled 'digest review,' not titled 'AI.' Rio has the folder link. Chris has the hop cap and the no-Jira line. You file the score table next to charter v0, the envelope, and the five-row interface table. Week 1 bought a map, a study OS, a mechanical picture, a vendor taxonomy, a prompt craft, a cost envelope, four interfaces, and this artifact. It did not buy a job offer or the right to say you know AI. Say instead: you can run a delivery prompt library, envelope a copilot, and refuse a chatbot-of-everything with a better interface. Specific is employable. Then you close the laptop at two hours. The people who wait for pretty will open day 8 with untitled chats. That is a browser history. You already decided artifacts beat badges. You beat them this morning, in a cluster of desks, with a deputy who almost pasted a claimant.

Diagram

Monday lab beats: ship the library

Repeats until the stop condition
01

Script + data rule

Standardizing the brief. Consumer tab closed. Rio hears forbidden data twice.

02

Show the folder

Five headers. Drafts do not feed the pack. Invented mess, not tidy secrets.

03

Happy + hostile

Digest holds amber on paint-us-green. RAID miner drops a false Pat owner.

04

Score and bump

Yes/no table. v1.1 changelog. Known failures in the header.

05

Note + freeze

Operating note on top. Deputy can find v1 in ten seconds. Thursday review booked.

06

Proof line

Positioning v0 gets the library. Kill if Thursday does not save twenty minutes.

last step feeds the first

Script, not an announcement. Sanitize, then run. Score yes/no. Freeze what survived. Operating note on top. Stop at two hours.

Practice

Ship the library

95 minutes

Five prompts + one operating page. This is artifact #1.

  1. Write all five with the header template (name, version, owner, date, model class, temperature, do-not-use-for, adversarial tests, changelog).
  2. Run each on a happy-path input and a hostile / empty input. Capture format / evidence / usefulness as yes/no in a table.
  3. Fix the worst two failures (usually: invented owners, no evidence, too long, sycophantic green).
  4. Write the operating note: allowed data, forbidden data, human review, storage place, change control, what to do when it invents.
  5. Add one proof line to your day-1 positioning v0. Tick the 'Delivery prompt library' artifact on the Capstone page.

Done looks like: A document (notes in this app are enough) with five versioned prompts, a ten-row score table, and a one-page operating rule. Not a folder of untitled chats.

Check yourself

Attempt in your notes first. Reveal is for after, not during.

  • What makes a prompt an operating asset instead of a personal trick?

  • What data rule belongs in every operating note?

  • What did week 1 actually buy you?

  • How do you test without a platform?

  • Name the five prompts in the library.

  • What is freeze in the prompt lifecycle?

  • What do you offer on Monday?

Terms from this day

Prompt library
A versioned set of prompts with owners, tests, and usage rules — like a process pack, not a chat history.
Sanitization
Stripping names, IDs, and confidential facts before using a tool that is not approved for that data.
Baby eval
A small, honest scorecard (format, evidence, usefulness) you run by hand before you have a harness.
Change control
Who may edit the prompt, and how you know production still matches the tested version.
Freeze
Declaring a prompt version as the one that may run; further edits require a new version and a re-test.
Operating note
The one-page runbook for the library: data rules, HITL, storage, owners, what to do when it invents.

Your notes for day 7

Saved on this device. Use this as the start of the artifact.