Your Library

My Prompts

Prompts you have saved or written yourself

Add a custom prompt:

Prompt sequences

Playbooks

Pick a situation. Follow the steps. Every step opens the right prompt.

The AI-Powered CSM

Courses

Start with Foundation. It is the course. Ten modules, about two and a half hours total, designed so you can be functional with AI in your CSM week within seven days. Advanced and Leaders are optional next steps, for if you want to go deeper or lead a team. Every course is free, runs in your browser, and gives you a verifiable certificate when you pass.

No account, no tracking, nothing leaves your machine. The trade-off: your progress lives in this browser, so if you switch device or clear your browsing data, it will not carry over. For a longer course, stick to the same browser to keep your place.

01
Start here

The AI-Powered CSM

The complete course. Ten modules of how to use AI well and safely in your day-to-day CS work, from prompts to call prep to churn risk. If you only do one course, do this one. About 2.5 hours total, designed so you are functional within a week.

10 modules · 2.5hComplete on its ownVerifiable certificate

Who it is for

Any CSM who wants to get genuinely good with AI without losing the human side of the job. No experience needed.

What you will be able to do

  • Turn a 45-minute call prep into a few minutes, without losing the depth
  • Write a churn risk read that catches what the data was quietly showing
  • Draft a QBR narrative or renewal case that sounds like you, not a robot
  • Spot expansion openings you would otherwise miss
  • Know exactly what is safe to put into AI, and what must never go in
  • Build a simple weekly habit so this sticks, instead of fading after a week

How it works

Ten short modules you do at your own pace. Each one teaches a real CS skill and ends with a knowledge check. Pass the checks and you earn the certificate. Everything uses real prompts you can take straight into your work.

Not started
Start here
The certificate you earn
Sample

Awarded on completion, with your name, a unique ID, and a verification line. Built to share on LinkedIn.

02
Optional · after Foundation

Advanced AI for CS

Optional next step. The builder's course, for CSMs who finished Foundation and want to go deeper. Design a version of AI that works the way you do, build custom assistants and a scoped agent, and become the AI authority in your team.

12 modulesAfter FoundationVerifiable certificate

Who it is for

A CSM who has the basics and wants to operate years ahead of their peers. Keen and curious, but not a developer. Everything is no-code or low-code.

What you will build

  • A digital twin: an AI that thinks and works like you
  • Custom assistants loaded with your own knowledge
  • A proper, tested prompt library
  • A real, working CS agent, end to end
  • A shared knowledge base that makes you your company's AI authority
Not started
Start the Advanced course
The certificate you earn
Sample

Awarded at 85% on all twelve modules. Your name, a unique ID, and the competencies you have proven.

03
Optional · if you lead a team

AI for CS Leaders

Optional next step. A different job. For team leads and CS managers who finished Foundation. How to take a whole team from scattered, nervous AI use to a capability that is consistent, safe, and measurable, and prove the value upward.

9 modulesFor team leadsAfter Foundation

Who it is for

Team leads, CS managers, and heads of CS. People responsible for others, not just themselves. You do not need to be technical.

What you learn

  • Why leading AI is a different job from using it
  • Reading your team's real use, including the shadow tools
  • Drawing a data line you can actually defend
  • Getting the whole team on board, not just the keen few
  • Proving the value upward, and being ready for the governance question
Not started
Start the Leaders course
The certificate you earn
Sample

Awarded at 85% on all nine modules. A credential that says you can lead a team into AI, safely and measurably.

Built for CSMs by a CSM

Become AI fluent without losing the human element.

Each module puts AI to work on your real accounts, not hypotheticals. Work through at your own pace and use the prompts live from day one.

AI leverage
Human judgement
0/10

Modules complete

No account, no tracking, nothing leaves your machine. The trade-off: your progress lives in this browser, so if you switch device or clear your browsing data, it will not carry over. For a longer course, stick to the same browser to keep your place.

Tier 1 · Foundations

The mindset, the method, the safety

M01–M04 · the four modules every CSM needs before anything else

MODULE 01

The AI mindset for CSMs

Why AI multiplies your best work rather than replacing it, and where it fits in your day

MODULE 02

Prompt engineering for CS work

The RCCO framework, and why CSM prompts fail when generic prompts succeed

MODULE 03

Choosing the right tool

Claude, ChatGPT, Copilot M365, Gemini, what each is genuinely best at for CS work

MODULE 04

Data hygiene and AI safety

What never to paste, how to anonymise, and how to stay on the right side of policy

Tier 2 · Core workflows

AI applied to the actual job

M05–M08 · call prep, communication, churn risk, expansion

MODULE 05

Account intelligence and call prep

Pre-call briefs, stakeholder maps, and ticket analysis, at ten times the speed

MODULE 06

Communication that lands

QBRs, renewal narratives, risk escalations, and editing the AI voice out

MODULE 07

Churn risk and health scoring

Four signal families and a dual-mode assessment, single account and full portfolio

MODULE 08

Expansion and whitespace

Product gap mapping, expansion signals, and business cases the buyer can forward

Tier 3 · Mastery

Make the practice yours

M09–M10 · build your operating system, hold the human core

MODULE 09

Building your AI operating system

From one-off prompts to a personal library and a weekly intelligence cadence

MODULE 10

The human element

When not to use AI, how to stay trusted, and turning fluency into career capital

Tier 1 · Foundation · Module 01

The AI mindset for CSMs

Why AI multiplies your best work rather than replacing it, and where it fits in your day

16 min 85% to pass

The lesson

75%

Roughly three quarters of a CSM's week goes on producing the inputs and outputs around conversations, the briefs, decks, emails, and reporting. Only about a quarter is spent in the live customer conversations where renewals are actually saved. This module is about flipping more of your week back to the part that matters.

1
Where your week actually goes

Before changing how you work, you need an honest picture of the work. Ask enough CSMs and the weekly split looks remarkably consistent: roughly a quarter of the week in live customer conversations, and the remaining three quarters spent producing the inputs and outputs around those conversations, briefs, summaries, decks, emails, CRM updates, internal reporting, inbox triage.

Here is the uncomfortable part: the live conversations are where renewals are saved and expansions are opened, and they are the part that gets squeezed. When an escalation lands, the production work does not shrink, the strategic work does. You walk in lighter, you react instead of lead. Every CSM knows this trade. Most have stopped noticing they make it daily.

The point

AI fluency, properly understood, is the discipline of moving hours out of the production column and into the conversation column. Not by cutting corners on the production work, but by changing your role in it.

Interactive tool
Your time-recovery calculator

Before reading another word, get your own number. Step 1. How many hours does a normal week cost you TODAY, working the way you work now, without AI? Honest averages, not worst weeks:

Step 2. How fluent will you get? (You can be honest; Module 02 onwards does the heavy lifting.)

Hold that number. The rest of this module, and honestly the rest of this course, is the technique for collecting it.

2
The junior analyst model

The single most useful mental model for working with AI: you have just been assigned a brilliant, tireless junior analyst. They have read essentially everything ever published. They write fluently in any style. They never get bored, never push back on tedious work, and turn drafts around in seconds.

They also have three defining limitations. They know nothing about your accounts until you brief them. They have no stake in being right, they will produce a confident answer whether or not the facts support one. And they carry no accountability, when the work goes out, your name is on it, not theirs.

Every good AI behaviour falls out of this model. You would not hand a junior analyst a task with no context and expect strategy, so brief properly (Module 02). You would not send their first draft unread, so edit and verify. You would not let them attend the negotiation for you, so keep the human work human (Module 10). And you would not re-explain the same task weekly, so build reusable briefings, which is all a prompt library is (Module 09).

The working relationship
YOU

Brief

Role, context, constraints, output. Everything it cannot know.

AI

Assemble

Synthesise volume into meaning. Draft. Surface patterns. In seconds.

YOU

Verify, edit, own

Read every line. Catch the inventions. Ship it under your name.

Your value moves from producing the work to directing and owning it.

The shift

Your value moves from producing the work to directing and editing it, which is exactly what senior people already do. Look at what your VP produces directly in a week: very little. Their output is judgement applied to other people's work. AI gives every CSM a production team. The ones who thrive learn to manage it like one.

3
What AI is genuinely good at, in CS terms

Skip the abstract capability talk. In CSM work specifically, AI excels at four families of task. You should recognise your week in each.

🧩
Synthesis

Collapsing volume into meaning. Fifty tickets into five themes. A year of call notes into a relationship arc. A 40-page contract into the six clauses that matter. The biggest hours live here, because synthesis is most of what "preparation" is.

✍️
Structured drafting

A competent first version of anything with a known shape: the QBR narrative, the escalation email, the exec summary, the follow-up. The shape is the prompt. The polish is yours.

🔎
Pattern surfacing

Noticing what is spread across too much material for one person to hold: the same complaint in three accounts, engagement decaying in a familiar pre-churn shape, a stakeholder going quiet.

🎭
Rehearsal

Playing the sceptical CFO, the procurement lead, the frustrated IT director, so the first time you hear the hardest objection is not in the room.

4
The four failure modes, and the catch for each

Fluency is as much about knowing where the tool breaks as where it shines. Four failure modes account for nearly every AI embarrassment in professional use.

⚠️
Confident invention

Given gaps, the model fills them, smoothly. The catch: constrain it ("do not invent, list missing information as open questions") and verify every fact you did not supply yourself.

🫥
Context blindness

It cannot know your champion is leaving, or that the CFO hates the word "partnership". The catch: the briefing is your job. Thin briefing in, generic output out, every time.

📋
Generic output

Ask for "a good renewal email" and you get the email everyone gets. The catch: specificity in, specificity out. The difference is almost always in the context, not the model.

🚫
Zero accountability

The model will not sit in the post-mortem. The catch: a rule you will meet again in Module 10, never send anything you could not defend line by line if asked "did you write this?"

Part 5 of 5
Worked example, the same task, three levels of fluency

The task: your manager asks for a quick health summary of an account before an internal pipeline review, due in an hour.

Level 0, no AI

You spend 45 minutes re-reading notes and tickets, write eight bullets from memory, and miss the pattern in the support data because there wasn't time to look. Adequate. Expensive.

Level 1, naive AI use

You type "write a customer health summary for a software account" and get 400 words of confident, generic filler that could describe any account on earth. You spend 20 minutes rewriting it and trust it less than your own bullets. This experience, not AI's actual ceiling, is why most CSMs quietly give up.

Actual output · the naive promptGeneric
Customer Health Summary Overall, the account remains in a generally stable position. Engagement levels are positive and the customer continues to derive value from the platform. There are some areas that may benefit from additional attention, and proactive outreach is recommended to strengthen the relationship. Key recommendations: • Continue regular check-ins to maintain alignment • Monitor usage trends for potential risks • Explore opportunities to demonstrate additional value

Read it twice: not one sentence is about this account. Every line is true of every customer that has ever existed, which means none of it is information.

Level 2, fluent use

You paste your raw material, last three call summaries, the ticket export, the usage trend, into a briefed, constrained prompt (you'll build it in Module 02). Ninety seconds later you have a structured draft: status, movement, top risk, top opportunity, recommended action. You correct one mischaracterised incident, add the political context only you know, and send it. Twelve minutes, end to end, and it's sharper than the Level 0 version because the synthesis actually covered everything.

Actual output · the briefed, constrained promptSpecific
STATUS: Stable but cooling. Renewal in 7 months. MOVEMENT: Usage flat for 2 quarters after 18 months of growth; CONTACT_1_CHAMPION reply times stretched from same-day to 4 days since the March reorganisation. TOP RISK: Convergence of relationship + commercial signals: champion latency and two early invoice queries from finance in 5 weeks. Innocent explanation exists (their finance migration), unverified. TOP OPPORTUNITY: 11 new users in the Madrid office (no prior usage); matches their announced European rollout. RECOMMENDED ACTION: Value-led re-engagement with CONTACT_1_CHAMPION this week; ask finance directly whether the queries relate to the migration. OPEN QUESTIONS: No data on executive sponsor engagement since January; Q1 NPS not provided.

The gap between Level 1 and Level 2 is not talent and it is not the tool. It is technique, briefing, constraining, editing, and technique is learnable. That is the rest of this course.

You will leave able to

  • Audit your own week and identify the 5 highest-leverage tasks to move to AI assistance
  • Explain the "junior analyst" model to a sceptical colleague in under a minute
  • Name the four failure modes of AI output and how to catch each one

Hands-on exercise

Part A, the audit. Track one full week in five categories: call prep, written comms, analysis, admin, live conversations. Mark each task H (human-essential) or A (AI-assistable). Most CSMs find 12+ hours of A-tasks. That's the time you can win back, write the number down, you'll use it in Module 10's ROI story.

Part B, the level test. Take one real A-task from your list and run the worked example yourself: do it manually (Level 0), then with a one-line prompt (Level 1). Keep both outputs. After Module 02, you'll run it at Level 2 and compare all three, your own before/after evidence, on your own work.

The human element: AI can summarise a customer's words. Only you can hear what they didn't say. The skill of reading a room, a pause, or a CC line stays yours, and becomes more valuable as the routine work disappears.

If you remember three things

  1. AI is a junior analyst: brilliant at assembly, zero context, zero accountability. Brief it like one.
  2. Four failure modes, confident invention, generic output, context blindness, zero accountability, and a catch for each.
  3. Specificity in, specificity out. Your weekly audit tells you exactly where the recoverable hours are.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

5 questions · 85% to pass

1. You ask AI to summarise a renewal call. The summary reads well and includes a line saying the customer confirmed budget for the expansion. Nobody said that on the call. What kind of failure is this?

Plausible content that was never in the source is confident invention, the most dangerous failure because it reads just like the true lines around it. The fix is a constraint ("use only what I've provided, list anything missing"), not more context. More context narrows the gaps but never closes them.

2. A CSM has used AI for six months. Their renewal rate is up. But they are slower on unscripted calls, lean heavily on their prep notes, and struggle when a conversation goes off-script. What is happening?

When AI does your thinking before calls but not during them, the live muscles waste. A higher renewal rate can hide declining depth and adaptability. This is the hardest failure to catch, because the numbers argue against you.

3. Your AI pre-call brief says the CTO has always been a strong supporter. You remember a conversation where they were clearly lukewarm. What do you do?

AI works from the data you gave it, which may be incomplete or read selectively. When your memory conflicts with the brief, that conflict is the important information, not noise. Verify it. Your instinct is also data, and often the most current you have.

4. In a sensitive renewal conversation, the customer asks: "Did you write this, or did AI?" What is the right answer?

Honesty plus ownership is the only answer that builds trust over time. Evasion costs you when it comes out later, and it always does. The test behind the answer: if you can't say "I reviewed it and stand behind every word" truthfully, you shouldn't have sent it.

5. AI handles the production work: drafting, summarising, analysing. So what is the part of your job that becomes more valuable, not less?

When the production work is cheap, the scarce thing is judgment: knowing which signal matters, reading what a customer isn't saying, and owning the decision. That is what AI can't do and what the rest of this course builds. Speed and volume are by-products, not the point.

Score: 0/5 ·

Tier 1 · Foundation · Module 02

Prompt engineering for CS work

The RCCO framework, and why CSM prompts fail when generic prompts succeed

16 min 85% to pass

The lesson

R C C O

Four parts to every good prompt: Role, Context, Constraints, Output. Get these right and the model writes like someone who knows your account. Skip them and you get the email that works for any account, which means it works for none.

1
Why CSM prompts fail when generic prompts succeed

Ask any model for "a poem about the sea" and you will get something decent: average is fine for generic tasks. Ask for "a renewal email" and you will get something decent looking, and that is the trap. An email that would work for any account is precisely an email that works for none. CS work does not tolerate average, because the entire value of a CSM's communication is its specificity to this customer, this history, this moment.

The specificity asymmetry

The model's output can only ever be as specific as your input. "Write a QBR summary" gives you the average of every QBR ever written. The same prompt plus your account's real year, their objectives in their own words, and the sensitive incident from March starts to sound like someone who knows the account. Because functionally it was: you briefed it.

So the discipline of this module is not "prompt engineering" in the hacker sense of magic words. It is briefing: the same skill you would use handing work to the junior analyst from Module 01. Four components, every time.

2
RCCO: the four components
🎩
Role

"You are a senior CSM preparing for a renewal call with a risk-flagged account." One line sets the vocabulary, the depth, and the relevance filter: what the model treats as worth mentioning. Match the role to the work, do not flatter the model.

📂
Context

Where CSM prompts live or die. The model knows nothing about your accounts. Every fact it does not have, it omits or invents. The call notes, the ticket export, their objectives, the thing that went wrong in March. This is where two CSMs using the same template get wildly different results.

🚧
Constraints

The guard rails. The single highest-value constraint in all of CS prompting: "Do not invent any facts, list missing information as open questions." Plus length limits, tone boundaries, UK English, and format rules.

📐
Output

The shape of the deliverable: "a one-page brief with these six sections", "a table with action, owner, date", "two drafts". It makes the result instantly usable and trains your eye to scan it in seconds.

The test before you send any context

Could a smart stranger complete this task with only what I have written here? If not, neither can the model.

Diagnosis becomes mechanical once you think in components. Output too shallow? Role or Context. Output confidently wrong? Constraints. Output rambling or unusable? Output spec. Output generic? Context, almost always Context.

3
Iteration: the first answer is a first draft

The single biggest behavioural difference between a Level 1 and a fluent AI user is what they do after the first response. The Level 1 user judges it: good or bad, keep or abandon. The fluent user treats it as the junior analyst's first draft and starts directing. Three follow-ups carry most of the value.

🕵️
"What did you assume that I did not tell you?"

The model fills gaps silently. This makes the filling visible. Run it on anything important and you will catch two or three assumptions you would never have spotted, one of which is usually wrong.

✂️
"Make this 50% shorter without losing the risk signals"

Models pad by default. Naming what must survive the cut gets compression without lobotomy. Senior readers get the short version. You keep the long one.

⚔️
"Now argue the opposite case"

The cheapest red team you will ever run. Think the account is safe? Ask for the case that it is churning. Somewhere in that argument is the objection the customer's CFO was always going to raise.

The point

Iteration is also where prompts become assets. When a follow-up fixes a recurring weakness, fold it back into the original prompt as a constraint. Your prompts should be evolving documents. Module 09 turns that habit into a library.

4
Reasoning you can audit

For anything analytical, churn assessment, negotiation prep, prioritisation, add one instruction: "Think step by step, showing your reasoning before your conclusions." Two effects. First, the quality of the conclusion measurably improves when the model works through the logic explicitly rather than jumping to an answer. Second, and more important for professional use: you can audit it. A bare conclusion ("this account is high risk") is take-it-or-leave-it. A reasoned chain can be checked line by line, and a wrong assumption gets caught in the logic before it becomes a wrong recommendation in front of your leadership.

Pair it with the disconfirming-evidence instruction you will meet again in Module 07: "What in this data argues against your conclusion?" You usually run an analysis because you already suspect the answer, which means you are primed to accept confirmation. Forcing the counter-case is the bias correction, built directly into the prompt.

One honest boundary

Visible reasoning is the model's account of its process, not a guarantee of truth. Audit it the way you would a junior analyst's logic: check the facts you supplied are used correctly and the inferences hold. Articulate does not mean accurate. The four failure modes from Module 01 still apply, reasoning just makes them findable.

Practice check · not scored
Diagnose the prompt
Prompt: "You are a senior CSM. Write a QBR summary for my account. Make it professional and thorough. Use clear sections."

Output: well-organised, confident, and padded with industry-generic claims and a made-up adoption statistic.

Which ONE change improves this prompt most?

Two components failed: Context (nothing real to work with) and Constraints (nothing forbidding invention), and they fail together, because an unbriefed model fills the vacuum. Role inflation and output formatting polish the container, not the content. And a model can't "double-check" a statistic it invented; the constraint has to prevent it, not audit it.
Part 5 of 5
Worked example: one task, four layers

The task: your champion at a key account emails that their new CFO is questioning the renewal cost. You need a reply that steadies the situation. Watch the output change as each RCCO layer is added.

No structure: "Write a reply to a customer questioning renewal cost"

A polite, generic defence of value that any vendor could send any customer. Mentions "partnership" twice. Your champion forwards it to the CFO and the CFO's scepticism is confirmed: this vendor is template city.

+ Role: "You are a senior CSM replying to a trusted champion whose new CFO is challenging renewal cost"

Better posture immediately: the reply addresses the champion as an ally to be armed, not a critic to be defended against. Still generic on substance, because the model knows nothing about the account.

+ Context: the products, the year's outcomes with numbers, the CFO's background, what the champion said exactly

Now it's recognisably about this account: it cites the actual outcomes, mirrors the champion's vocabulary, and anticipates a finance-shaped objection because you mentioned the CFO came from a cost-cutting mandate.

+ Constraints and Output: "Do not invent figures. Under 150 words. Give the champion two forwardable sentences for the CFO. No 'partnership'. UK English."

The deliverable: a tight reply your champion can act on in sixty seconds, two sentences engineered to be forwarded upward, every number traceable to what you supplied. Total time, including gathering context: about eight minutes. The Level 0 version of this email historically took you forty, and didn't include the forwardable sentences, because under time pressure nobody thinks of the clever extra.

That's the module. The RCCO builder below assembles the skeleton for you; the exercise makes it muscle memory. From here on, every vault prompt you open is this framework wearing work clothes.

Interactive tool
RCCO prompt builder

Build a structured prompt from the four components, then copy it straight into your AI tool.

You will leave able to

  • Write an RCCO-structured prompt from scratch for any CS task
  • Diagnose why a prompt produced weak output and fix the right component
  • Use reasoning chains and counter-argument follow-ups to pressure-test AI analysis

Hands-on exercise

Take a real email you sent last week. Write an RCCO prompt to reproduce it. Compare the AI version with your original, then merge the best of both. Most people find the AI version is better structured and theirs is better judged. That's the whole course in one exercise.

The human element: the Context component is irreplicable expertise. Two CSMs with the same prompt template get wildly different results because one knows the account and one doesn't. Your context is your moat.

Includes vault prompts: Pre-call intelligence brief · Exec summary compressor

If you remember three things

  1. RCCO every time: Role, Context, Constraints, Output. Diagnosis is mechanical once you think in components.
  2. The highest-value constraint in CS: “do not invent facts; list missing information as open questions.”
  3. The first answer is a first draft: surface the assumptions, compress it, then make it argue the opposite case.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. You run a renewal risk prompt using RCCO. The output includes specific competitor win-rate statistics you never supplied. Which RCCO component failed, and what is the fix?

Invention is always a constraints failure. "Do not state anything not present in the notes, list missing information as open questions instead" is the single highest-value line in CS prompting. It converts confident fiction into an honest gap list you can actually act on.

2. "You are an assistant" vs "You are a senior CSM with ten years in complex B2B accounts, preparing for a renewal at risk." What is the most accurate description of what changes?

Role framing shifts the entire distribution of the response, what the model treats as relevant, what it surfaces first, what it assumes you already know. A "senior CSM" frame surfaces risk, politics, and relationship complexity. No frame surfaces generalities that apply to every customer everywhere.

3. You have a solid first draft of a renewal risk email. The highest-value follow-up prompt is:

The adversarial variant reveals your draft's blind spots. You are sending into a specific political context, the version optimised for a receptive champion may fail completely with a resistant or sceptical one. Generating the harder version first forces you to confront that gap before it confronts you.

4. Why request step-by-step reasoning on a complex renewal risk assessment, rather than just asking for the conclusion?

Visible reasoning is an audit trail. A wrong assumption in step 2 of the logic shows up as a wrong step, not as a plausible-sounding conclusion that only falls apart in front of your VP. The chain is where you do quality control, not the summary.

Score: 0/4 ·

Tier 1 · Foundation · Module 03

Choosing the right tool

The routing mindset that works with any tool your organisation uses

12 min 85% to pass

The lesson

3questions

Being fluent is not about picking the one best tool. It is about routing: sending each job to the right tool in about five seconds. Three questions, in order, make that decision for you. Notice "which model is cleverest this month" is not one of them.

1
Fluency is routing, not loyalty

Ask ten CSMs which AI tool is best and you will get ten loyalty declarations: "I am a ChatGPT person", "our company is all-in on Copilot". Loyalty is the wrong frame. The fluent CSM is a router: different jobs go to different tools, and the routing decision takes about five seconds once you have internalised the question stack.

The cost of loyalty

Tool loyalty quietly costs you in one of two directions: either you are pasting customer data into a tool that should not have it, or you are accepting mediocre output from a tool that was never built for the task in front of you.

The question stack, in strict order: where must the data stay, what kind of job is this, and where does the work live. That is the whole decision.

2
The three routing questions
🔒
1. Where must the data stay?

This gate comes first because it is the only one you cannot edit your way out of. If the task touches customer data, the choice is between tools your organisation has approved for that data and everything else. The approved list is a fact held by IT, not a feeling. If you have not seen it, asking for it is this module's first action.

🧰
2. What kind of job is this?

Frontier assistants for deep reasoning and long-form quality, when the thinking is the product. Embedded copilots when the work happens inside an app you are already in. Vertical CS platforms that know your data model natively. Friction beats brilliance for in-flow work: an 85% reply inside the Outlook thread beats a 95% reply that needed three copy-pastes.

🔄
3. When the models change?

A new model ships roughly every quarter. Do not chase every release or ignore them. Keep your three most-used prompts with real anonymised inputs as a personal benchmark. When a big release lands, run them side by side and judge on faithfulness, edit distance, and reasoning quality.

The point

League tables measure averages. Your personal benchmark measures what actually matters to your specific work, and you do not do average work. Twenty minutes, once a quarter, and your whole library travels safely across each new model generation.

4
Anonymisation widens every option

One practical note that runs through all three questions: anonymisation (Module 04's placeholder scheme) widens your options dramatically. A churn analysis about CUSTOMER_A with CONTACT_1_IT can travel to whichever tool does the best analysis, because the sensitive layer never left your desk. The routing questions get easier the moment the data is no longer identifiable.

Part 5 of 5
Worked example: one Tuesday, four routing decisions
09:10 · A reply to a stakeholder email sitting in Outlook

Embedded. Copilot M365 drafts in-thread with the history right there; data never leaves the tenant; friction is zero. The frontier models would write a marginally better paragraph that costs ten times the workflow.

10:30 · 200 support tickets need a pattern analysis before a risk call

Frontier, anonymised. This is synthesis-heavy thinking-work, exactly what the strongest reasoning model you're approved for is for. Tickets exported, placeholder scheme applied, deep analysis back in minutes.

13:00 · "What's our renewal date and last QBR score for this account?"

Vertical tool. Your CS platform answers from the system of record in one query. Sending this to a frontier model means assembling context the CRM already has, to get an answer it already knows.

15:45 · The renewal narrative for your hardest account, CFO audience

Frontier, full ritual. Maximum stakes, thinking is the product: briefed prompt, disconfirming pass, two drafts, your judgement on every line. This is the 10% where the best available model earns its trip.

Four decisions, maybe twenty seconds of routing thought in total. That's the skill: not knowing every model's benchmark scores, but knowing your own question stack cold.

Practice check · not scored
Route the task
Your manager asks for a one-paragraph summary of this morning's Teams call with a customer, for the account channel, within the hour. The transcript sits in Teams.

Best route?

Question one: customer data, must stay in the tenant. Question two: routine synthesis, not deep reasoning. Question three: the work lives in Teams. All three point the same way, and the frontier detour would add friction, risk, and approximately nothing.
You suspect a strategic account is quietly churning and want the rigorous dual-mode risk assessment from Module 07, including a devil's-advocate pass on your own theory.

Best route?

This is the 10%: multi-signal reasoning, a disconfirming pass, an intervention plan. Embedded copilots and platform health scores will give you something; the strongest reasoning model you're allowed, fed anonymised context, gives you the analysis a renewal actually deserves. The platform score is an input to this work, not a substitute for it.

You will leave able to

  • Route any CS task to the right tool using the three-question framework
  • Run a personal benchmark to evaluate any new model in 15 minutes
  • Explain to your manager or IT team why you use the tools you use

Hands-on exercise

Run the same vault prompt (pre-call brief) on two tools you have access to. Score each output on accuracy, structure, tone, and usefulness out of 10. Keep the scorecard, it's the start of your personal benchmark.

The human element: tool choice is also a trust signal. Using the compliant tool for customer data, even when a better one exists, is the kind of judgement that makes leadership comfortable scaling AI across the team. Be the CSM who gets that right.

If you remember three things

  1. Route, don't pledge loyalty: where must the data stay → what kind of job is this → where does the work live.
  2. Friction beats brilliance for in-flow tasks; save the frontier trip for work where thinking is the product.
  3. Own a personal benchmark: your three most-used prompts, re-run on every major release, twenty minutes a quarter.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. You are drafting a sensitive renewal update inside an existing Outlook thread. You have access to both a frontier AI model in a browser tab and Microsoft Copilot embedded in Outlook. What is the most important factor in your tool decision?

Compliance routes before quality. If customer data must stay in your Microsoft tenant, Copilot is the answer regardless of which model writes prettier prose. Getting this order of operations right is what gives you the licence to use AI at all, and what protects you when someone asks which tool you used.

A major new AI model ships with significant benchmarks. What is the right first move for a fluent CSM?

Fifteen minutes, your own benchmark. Static league tables go stale within weeks; your own test against real work from your actual role never does. The fluent CSM evaluates tools the same way they evaluate vendor claims, against their own evidence, not someone else's.

3. For post-call notes you have two options: Tool A, a frontier model in a browser tab that produces noticeably better prose; Tool B, embedded in your CRM with one-click access. Which do you choose, and why?

The embedded tool you use every single call beats the better writer you open a tab for on two calls a week. Habit compounds. The tool with the lowest friction between intent and action will define your actual workflow, not the tool with the highest theoretical ceiling.

4. Your company has approved one AI tool for customer data. A colleague argues that a different tool writes significantly better renewal emails. What do you do?

Visibly choosing the approved tool is exactly what makes leadership comfortable scaling AI to the whole team. Rogue tool use, even with good intentions and anonymised data, creates the incidents that restrict AI for everyone, including you. Module 04 covers when anonymisation legitimately opens additional options.

Score: 0/4 ·

Tier 1 · Foundation · Module 04

Data hygiene and AI safety

What never to paste, how to anonymise, and how to stay on the right side of policy

14 min 85% to pass

The lesson

100vs1

A hundred brilliant AI-assisted briefs build your reputation slowly. One customer name in the wrong tool can undo all of it in an afternoon. This is the asymmetric risk that decides whether you get to keep using everything else in this course.

1
The asymmetry that rules everything

Data safety is not the boring compliance chapter of AI fluency. It is the asymmetric risk that decides whether you get to keep practising everything else. That is why this module sits at the end of Foundation, before a single exercise touches real account data.

Two surfaces, that is all

There are exactly two places things go wrong: what goes in (the data you paste, upload, or reference) and what comes out under your name (the drafts you send, the facts you repeat, the commitments you did not notice). Neither is difficult. Both have to become reflex.

2
What goes in: the four data classes

Before anything touches a model, classify it. Four classes cover everything a CSM handles.

👤
Personal data

Names, emails, job titles, opinions, anything identifying a real person. Under GDPR, pasting it into an AI tool is processing personal data, full stop. "The vendor does not train on my data" answers a different question. What matters is the data processing agreement and the approved list.

💼
Commercially sensitive

Pricing strategy, discount floors, legal positions, anything under NDA, unannounced plans. One place only: an approved enterprise tool, anonymised, with explicit clearance. Never a personal account. The convenience is never worth the legal risk.

🔑
Security material

Credentials, API keys, architecture diagrams, security questionnaire answers. Never. There is no anonymised version of a password.

Everything else, after prep

Ticket themes, usage patterns, meeting structures, your own drafts. The vast majority of CS work. Once it has been through the placeholder pass, it can travel to whichever approved tool does the job best.

The gate in front of all four

The approved-tools list: a real, named list held by your IT or security team. If you have not seen it, requesting it is today's action. The free tier of anything, and any personal account, is off-list by definition.

3
The placeholder discipline

Anonymisation done badly means deleting things, which destroys the analysis: a churn assessment where every actor is "[redacted]" cannot reason about who influences whom. Anonymisation done well means consistent placeholders: CUSTOMER_A for the account, CONTACT_1_IT and CONTACT_2_FINANCE for people (the role suffix keeps the politics), COMPETITOR_X for rivals. The model reasons about the relationships perfectly, because the relationships survived. Only the identities stayed home.

Three habits that make it work

Keep a key document (placeholders to real names) in a local file that never goes near an AI tool. Substitute before the text enters the prompt, find-and-replace takes ninety seconds. And watch the three leaks people forget: transcripts (full of names), exports (columns you did not scroll to, email signatures), and screenshots (tab titles and taskbar around the thing you meant to share).

The reflex to build: the ten-second check. Would you be comfortable if this exact text appeared in an email to the customer's legal team? Hesitation is data.

4
What comes out: owning the output

The second surface gets less attention and bites just as hard. Three disciplines.

🔍
Verify before it ships

Every fact, figure, date, and name in an AI draft gets checked against the source before it leaves your hands. Confident invention does not announce itself, it sits in fluent sentences next to true ones.

🤝
Hunt the hidden commitments

AI drafts love generous closing lines: "we will have that to you by Friday", "happy to include that at no extra cost". Read every draft for promises you did not authorise before it goes out.

✒️
Own every word

Your name is on it. The standard from Module 01, and again in Module 10: never send anything you could not defend line by line if asked "did you write this?"

Part 5 of 5
Worked example: a sensitive analysis, end to end

The task: a churn-risk analysis on a strategic account, using a 180-row ticket export and your call notes. Watch the discipline applied.

The raw material fails the check

The export contains four named contacts, two email signatures with mobile numbers, and one ticket where your champion criticises her own CIO by name. The call notes mention the customer's unannounced restructure. Pasting this anywhere, even a approved tool, fails the ten-second check on the restructure alone.

The placeholder pass, ninety seconds

Find-and-replace in a text editor: the account becomes CUSTOMER_A, the four contacts become CONTACT_1_CHAMPION, CONTACT_2_CIO, CONTACT_3_IT, CONTACT_4_PROCUREMENT. Signatures and numbers deleted (they carry no analytical value). The restructure stays in, described as "an unannounced internal reorganisation", because the analysis needs it but the specifics don't travel. Key document saved locally.

The run, and the output discipline

The prepared text goes to your strongest approved tool with the Module 07 prompt. The analysis comes back sharp, the politics intact: "CONTACT_1_CHAMPION's criticism of CONTACT_2_CIO suggests the relationship risk sits above your champion, not below." You verify the two statistics it cites against the export (both real), translate the placeholders back using your key, and brief your manager. Total overhead for complete safety: about three minutes.

Three minutes. That's the entire cost of doing this professionally, and it's the difference between AI fluency as a career asset and AI fluency as the subject of a very uncomfortable meeting.

Practice check · not scored
Classify before you paste
A colleague pastes a full ticket export, real names included, into a free consumer AI tool, reasoning: "It's fine, this one doesn't train on your data."

What's the flaw in that reasoning?

Two independent failures: personal data was processed in a tool with no data processing agreement, and the tool isn't approved. The no-training policy is real but answers a question nobody asked. The fix costs ninety seconds: placeholder pass, then a approved tool.
Three items on your desk: (a) a ticket-theme analysis for a strategic account, (b) your discount floor for the upcoming renewal negotiation, (c) a Teams transcript you want summarised.

Which routing is right?

(b) is commercially sensitive strategy: there's no anonymised version of your own negotiating floor, because the number IS the secret. (a) is the everything-else class, safe after the placeholder pass. (c) is Module 03 routing: the embedded tool summarises in place and the data never moves. Blanket avoidance fails too: refusing all three just means doing the safe two slowly.

You will leave able to

  • Apply the 10-second pre-paste check automatically
  • Anonymise an account scenario in under two minutes without losing analytical value
  • State the three questions to ask IT or legal before adopting any AI tool

Hands-on exercise

Take a real (sensitive) account summary. Produce a fully anonymised version using the placeholder scheme. Run a churn-risk prompt on it and verify the analysis still works. This becomes your reusable anonymisation template.

The human element: data judgement cannot be delegated to the tool. The model will accept anything you paste. The standard is yours to hold, and holding it visibly is what earns you the licence to go further with AI than anyone else on your team.

If you remember three things

  1. Classify before you paste: personal data, commercially sensitive, security material, or everything-else-after-prep.
  2. Placeholders preserve the politics (CONTACT_1_CHAMPION); the key document never goes near an AI tool.
  3. Everything ships under your name: verify the facts, hunt the hidden commitments, own every line.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. Before pasting content into an AI tool, which single question matters most?

One question, every paste, no exceptions. If the answer is no: strip it, generalise it, or move to the approved tool. The discipline takes ten seconds; the incident it prevents can take a year to resolve and permanently affects how much your organisation trusts AI in the hands of the CS team.

2. A colleague says anonymisation is "theatre", "the model does not know who this is anyway." What is fundamentally wrong with this reasoning?

Under GDPR and most enterprise data agreements, processing personal data means handling it, and transmission to a third-party server is handling, full stop. What happens afterwards (training, retention, deletion) is a separate legal question. The moment you send it is the moment that counts.

3. Your company has no AI usage policy yet. The most responsible approach is:

Existing data policies, confidentiality agreements, and customer contracts already constrain AI use even without an explicit policy. "No AI policy" does not mean "no applicable policy." And raising the question positions you as the responsible AI voice in the organisation, that is a career advantage, not a burden.

4. You find a prompt template in a shared team folder. It contains real customer contract values and health scores as example inputs. What do you do before using it?

Specific customer data in prompt examples is a data governance issue regardless of which tool the prompt is built for or where it is stored. A placeholder scheme, CUSTOMER_A, $XXXK ARR, HEALTH_SCORE_67, gives identical prompt engineering quality with zero data exposure. Takes two minutes and is the only version you should share.

Score: 0/4 ·

Tier 2 · Core workflows · Module 05

Account intelligence and call prep

Pre-call briefs, stakeholder maps, and ticket analysis, at ten times the speed

14 min 85% to pass

The lesson

40to15min

Five meaningful conversations a week, forty minutes of prep each. Call prep is one of the biggest line items in your calendar. AI compresses the gathering and synthesis to ten or fifteen minutes. The judgement about what matters stays exactly where it always was: with you.

1
The highest-frequency skill in the job

Five meaningful conversations a week, forty minutes of honest preparation each: call prep is quietly one of the biggest line items in your calendar. High frequency, high leverage, and the gap between prepared and unprepared is visible to the customer within ninety seconds.

The standard this module sets

Walk into every call knowing more than anyone expects you to. Not more data, more synthesis: what changed, what it means, and the one question that proves you have been paying attention.

2
The fixed skeleton

The pre-call brief in the vault uses the same six sections every single time: status line (the account in one sentence), since we last spoke (what changed, with dates), open items (theirs and ours, with owners), risk signals (anything that smells off, however faint), opportunity (anything that smells like expansion), and the one question (the single thing to ask that this customer would not expect a vendor to know to ask).

Why fixed matters

The skeleton being fixed is the entire trick. The fiftieth time your eyes land on a brief with identical structure, you are no longer reading it, you are scanning it, and deviations jump out the way a wrong note does in a familiar song. Fixed skeleton, fluent scanning, pattern recognition for free.

3
The layering technique

A good brief uses your internal data. An exceptional one adds two layers on top, and the third is where the magic happens.

Three layers, in order
3

The connection pass

"Connect what is happening in their world to what is happening in our account. What should I infer, and what should I ask?" The question a strategic adviser always asks.

▲ sits on
2

Public intelligence

Their results, leadership changes, product launches, sector news. Five minutes of searching, pasted beneath the base.

▲ sits on
1

The base brief

Call notes, tickets, usage, CRM history through the skeleton prompt. The baseline everyone can reach.

Each layer holds up the one above. The magic lives at the top, but only because the bottom two are there.

Where briefs become intelligence

Their CEO announced "ruthless cost discipline" in October and your expansion conversation stalled in November? Those are not two facts, they are one story. The CSM who walks in already understanding that story has a fundamentally different conversation from the one who asks "so, how is everything going?"

4
The two standing assets: the map and the tickets
🗺️
The stakeholder map

Plots everyone who matters by influence and sentiment, quarterly and after any reorg. Read it the uncomfortable way: the most dangerous output is the gaps, the influential roles where you have no relationship at all. Accounts are rarely lost to the sceptic you were managing, they are lost to the budget holder you had never met.

🎫
Ticket analysis

Separate routine friction from strategic risk. Fourteen password resets are friction. The signals hide in trend (volume bending upward), tone (frustration entering patient people's language), and seniority (a director personally filing tickets is not a ticket, it is a message). Tell the model to make this split, or it will summarise the password resets.

Part 5 of 5
Worked example: fifteen minutes before the check-in

The call: a monthly check-in with a strategic account, renewal in five months. Watch the layers go on.

Minutes 0 to 6, the base brief

Last two call notes, the ticket export, usage summary into the skeleton prompt (placeholders applied per Module 04). Out comes the structured brief. Status: stable but quiet. Since last spoke: usage flat, one new integration ticket. Risk signals: champion's reply times have stretched from same-day to four days. The skeleton surfaces that last one only because response patterns are in your notes.

Minutes 6 to 11, the public layer

A search on the account turns up their interim results from last week: revenue miss, and a new CFO starting next month, announced with a brief about "operational efficiency". Pasted under the brief.

Minutes 11 to 15, the connection pass

"Connect their world to our account." The output: a slowing champion plus an incoming efficiency-mandated CFO five months before renewal suggests your champion may be distracted by internal repositioning, and every supplier contract is about to get fresh scrutiny. Inferred risk: the renewal becomes a procurement event. The one question: "I saw the announcement about your new CFO; how is the team preparing, and what will that change for how decisions like ours get reviewed?" You walk in as the supplier who already understands their quarter. The unprepared version of you would have opened with the weather.

Practice check · not scored
Run the connection pass yourself
Base brief: usage stable, tickets routine, but your champion hasn't replied in three weeks. Public layer: the company has just announced a new CFO with a cost-reduction mandate.

What's the connection-pass reading?

The connection pass produces an inference and an action, not a panic and not a shrug. The two facts together suggest distraction plus incoming scrutiny: a re-engagement with a ready value case beats both the do-nothing read and the catastrophising one. "Champion is leaving" might be true, but nothing in the data says it yet; that's invention wearing an insight costume.
This month's ticket export: fourteen password resets, two API-timeout tickets from the integration lead, and one ticket from a VP of Operations asking how to bulk-export their data.

Which one is the strategic signal?

Seniority changed the channel: a VP personally filing a ticket is a message, and bulk data export is one of the classic pre-departure patterns (it's also, sometimes, perfectly innocent reporting). The discipline isn't to panic; it's to notice, investigate, and find out which one it is, this week. The password resets are friction; the timeouts matter to the integration lead and should be fixed, but neither moves the renewal.

You will leave able to

  • Produce a one-page pre-call brief from raw inputs in under 15 minutes
  • Build and maintain AI-assisted stakeholder maps that flag relationship gaps
  • Turn a support ticket export into a themed risk analysis for any account

Hands-on exercise

Pick your next real call. Run the full brief workflow, base brief, public intelligence layer, connection step. After the call, score the brief: what did it nail, what did it miss, what surprised you? Refine your version of the prompt accordingly.

The human element: the brief gets you to the starting line; the call is still yours. AI cannot build rapport, read hesitation, or know that your champion sounded flat last time. Prep faster precisely so you can be more present.

Includes vault prompts: Pre-call intelligence brief · Stakeholder map builder · Support ticket pattern analysis

If you remember three things

  1. A fixed brief skeleton turns reading into scanning, and anomalies jump out on their own.
  2. Layer it: base brief, public intelligence, then the connection pass, where summaries become intelligence.
  3. Ticket signals live in trend, tone, and seniority. A director filing a ticket is a message, not a ticket.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. Your AI pre-call brief contains four sections: recent support activity, stakeholder map, open commitments, and risk signals. Which element requires the most scrutiny before you act on it?

"The CTO has always been a strong supporter" is exactly the kind of claim AI generates smoothly and confidently from ambiguous notes. If your read of that CTO differs, that conflict is the important information, verify against CRM before any call where you intend to rely on it.

2. You use a fixed seven-section skeleton for every pre-call brief. A colleague says you should vary it account by account for freshness and relevance. Why is consistency the better choice?

The skeleton is for you, not the model. When risk signals are always in slot 5, a thin slot 5 is an immediate flag, before the call, not in it. Variation hides omissions. Consistency makes them impossible to miss, which is the whole point of a system.

3. Your ticket analysis for a large account concludes: "primary ongoing challenge is onboarding complexity." Onboarding completed six months ago. What most likely happened, and what is the fix?

Models analyse what you give them. A ticket export without date filtering mixes historical noise with current signal. The filtering step is mandatory before analysis, not optional. "Last 90 days only" is often the single highest-value constraint you can add to any ticket analysis prompt.

4. One hour before a QBR, your champion calls to say a new senior executive is joining who you have never spoken to. What is your move?

This is exactly what AI is built for, high-value preparation in compressed time. Three minutes with a good prompt gives you more than nothing. The executive's LinkedIn, their title, their function's likely concerns: that is enough to open well and avoid a misstep in the first five minutes.

Score: 0/4 ·

Tier 2 · Core workflows · Module 06

Communication that lands

QBRs, renewal narratives, risk escalations, and editing the AI voice out

14 min 85% to pass

The lesson

Voice+specifics

Communication is the surface your whole job is judged on. AI gives every CSM the same tireless first-drafter, so inboxes are filling with prose that is fluent, polished, and eerily identical. The writing that lands sounds like a specific human who knows this specific customer. AI supplies structure and speed. You supply voice and specifics.

1
The visible surface of the job

Most of your work is invisible to the people who decide your account's future. They do not see the triage, the internal escalations, the fifteen-minute briefs. They see what you send: the QBR, the renewal case, the email after the difficult call. Communication is not one CSM skill among many, it is the surface on which all the others are judged.

The gift and the trap

The gift: a tireless first-drafter. The trap: every CSM now has the same one, and inboxes are filling with prose that is fluent, polished, and eerily identical. Sameness is the new mediocrity. The whole module is one division of labour: AI supplies structure and speed, you supply voice and specifics.

2
Narrative first, slides second

The default QBR failure is template-first: open last quarter's deck, refresh the numbers, present data without meaning. The fix is to demand the story before the deck exists. The vault's QBR prompt asks for a narrative arc in prose: where this account started the quarter, what actually happened, what it means, and where we go next, with every claim tied to evidence. Only once that arc reads true do you ask for the slide structure.

The test

If you deleted every chart, would the meeting still have a point? When the answer is yes, the charts become evidence for an argument rather than a substitute for one, and the customer leaves remembering a story about their own progress, which is the only thing anyone ever remembers from a QBR.

3
The two-draft technique

Hard messages, risk escalations, unwelcome news, pushback on an unreasonable ask, sit somewhere between maximum clarity and maximum relationship care. The mistake is asking AI for the message. The technique is asking for two drafts at different points on that spectrum: one direct and unhedged, one cushioned and warm, both carrying the same facts and the same ask.

Why two

The spectrum point is not a writing decision, it is a relationship decision, and only you hold the relationship. The burned engineer needs the direct draft. The politically exposed sponsor needs the cushioned one. Two drafts cost the model nothing and hand the only judgement that matters back to the only person qualified to make it. One more habit: write for the forward. Assume your email gets sent upward with one line of commentary added.

4
The de-AI editing pass Updated June 2026

AI prose has tells, and your customers are learning to spot them at the same rate the models are getting subtler. The 2023-vintage tells (delve, leverage as a verb, "I hope this email finds you well") still appear, but the modern frontier models in 2026 produce most of their tells at the structural level rather than the word level. The pass below targets both. It takes three minutes.

✂️
Cut the word-level tells

"I hope this email finds you well." "Delve." "Leverage" as a verb. Less common now in frontier output but still survives in budget tools and default ChatGPT. If a phrase could open any email to any customer from any vendor, it goes.

🧱
Cut the structural tells

The symmetrical opening that names three things you'll cover. The rule-of-three sentence in every paragraph. "It is not just X, it is Y." "Ultimately, ..." as a conclusion. The closing paragraph that gives equal weight to two sides. These are what 2026 frontier models do almost without prompting, and they read as AI faster than any word does.

📌
Re-inject what only you know

The specific date, the colleague's actual name, the reference to what they said on Tuesday's call. One genuine specific is worth three paragraphs of polish, because specifics are the one thing the model cannot supply and the customer cannot miss.

🔊
Read it aloud

The fastest authenticity test there is: anywhere your actual voice would not say it, rewrite it until it would. Teach the model your voice up front by pasting a sample of your real writing, but still do the read-aloud pass at the end.

Part 5 of 5
Worked example: the escalation email, before and after

The situation: a data-sync failure on your side corrupted a week of the customer's reports. The fix is in, but trust is bruised, and the email matters.

The naive draft (one generic prompt, sent as-is)

"I hope this email finds you well. I wanted to reach out regarding the recent issue you may have experienced. We deeply value your partnership and are committed to delivering excellence..." Three paragraphs in, the incident still hasn't been named, no dates appear, and the apology could be from any vendor about anything. The customer reads it as exactly what it is: nobody home.

The technique applied

Two drafts requested with the full incident context: one maximally direct, one warm. This customer's sponsor is an ex-engineer who has been generous through the incident; the blend leans direct with one human line kept from the warm draft.

After the de-AI pass and the specifics

"Hi Sarah, the sync failure between the 4th and the 11th corrupted the weekly reports your team pulled in that window, and that's on us. Here's what happened, what we've fixed, and how we'll catch this class of fault before you ever see it again..." Names, dates, plain ownership, the precise mechanism in one paragraph, the prevention in another, and a closing line referencing the patience her team showed on Thursday's call. Ninety seconds longer to produce than the naive draft. A different universe to receive.

Practice check · not scored
Run the de-AI pass
AI draft, final paragraph: "We remain fully committed to your success and will continue to leverage our resources to ensure a seamless experience going forward. Please don't hesitate to reach out should you have any questions."

What does the de-AI pass do here?

Every phrase in that paragraph could close any email from any vendor to any customer, which means it communicates exactly nothing. Word-swaps polish the emptiness. The pass replaces it with the two things generic prose can never carry: a concrete next step with a date, and a line that proves a human who knows this account wrote it.
You must tell a customer their feature request, promised "consideration" by a colleague who has since left, is not on the roadmap. Their sponsor is a blunt-spoken CTO who has twice complained about vendors "wrapping bad news in marketing".

How does the two-draft technique resolve?

The spectrum point is a relationship decision, and this relationship has stated its preference twice. For this reader, directness IS the relationship care: name the broken promise, own it, state the reality, offer what's actually possible. The 50/50 default and the always-soften rule both ignore the only data that matters: who is reading.

You will leave able to

  • Build narrative-first QBRs that lead with meaning, not metrics
  • Produce renewal value cases written in the customer's own language
  • Run the two-draft technique for high-stakes messages, and the de-AI editing pass on everything

Hands-on exercise

Take your next real QBR. Run the narrative-arc prompt before building a single slide. Present the story to a colleague in 60 seconds. If they can repeat it back, build the deck. If not, iterate the narrative, not the slides.

The human element: AI gives you a hundred competent sentences. Choosing the one this customer needs to hear, and being accountable for it, is the job. Never send a high-stakes message you couldn't defend line by line.

Includes vault prompts: QBR narrative builder · Renewal value case · Risk escalation (two-draft) · Meeting notes to actions

If you remember three things

  1. The division of labour: AI supplies structure and speed; you supply voice and specifics. Neither is optional.
  2. Hard messages get two drafts, because the clarity-vs-care decision is a relationship decision only you can make.
  3. The de-AI pass: cut the tells, re-inject what only you know, read it aloud.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. Your AI-drafted QBR narrative opens with three paragraphs of industry context and product background before reaching the customer's results. What caused this, and what is the fix?

Fix: "The audience is the VP of Operations who has been a customer for two years. They know the product and our methodology. Skip all company and product introduction. Open with their outcome against the goal they set at the last QBR." That single instruction changes everything the model treats as relevant.

2. You need to escalate a churn risk to your VP. Which approach produces the most useful communication?

The model needs your analysis to help you communicate it. Without your diagnosis of cause, severity, and recommended action, it generates informed-sounding speculation, and informed-sounding speculation is the last thing you want in an executive escalation. Your VP is making a resource decision; give them what they need to make it.

3. What does the "de-AI" editing pass primarily target?

Cut the hedged openings, the triple adjectives, the "I hope this finds you well." Then add the detail only you could know: the specific thing the customer said in last week's call, the exact number from their own reporting, the context that shows you were paying attention. That is your voice back in the document, and customers can tell.

4. You are writing a renewal business case that your champion will forward to their CFO. The guiding principle is:

Your champion will not be in the room with their CFO. The document has to make the case itself, in language the CFO uses, against the metrics they track. If it requires your champion to translate, it will fail somewhere in that translation, and you will not be there to rescue it.

Score: 0/4 ·

Tier 2 · Core workflows · Module 07

Churn risk and health scoring

Four signal families and a dual-mode assessment, single account and full portfolio

14 min 85% to pass

The lesson

3quarters

No account leaves on the day it leaves. The decision was usually made quietly, two or three quarters earlier, with the signals sitting in your data the whole time. This module is about closing the gap between when a signal appears and when you act on it.

1
Churn is a process, not an event

No account leaves on the day it leaves. The decision was made quietly, internally, often two or three quarters earlier, and the signals were sitting in your data the whole time: a champion whose replies stretched, a procurement question that arrived early, a usage line that bent. Run a post-mortem on any lost account (the vault has the prompt) and the most painful column is always the same one: the gap between when each signal first appeared and when someone first acted on it. That gap is detection lag, and shrinking it is the entire job of this module.

The point

AI cannot tell you whether an account will churn. But it can read more signals, more consistently, more often than you can, so the signals reach your judgement while there is still time for it to matter.

2
The four signal families

Risk signals come in four families, and the families matter more than the individual signals. Watch them as a group, not a checklist.

📊
Engagement

Usage volume and breadth, login patterns, feature adoption, who has stopped showing up in the data.

🤝
Relationship

Champion behaviour: reply latency, meeting attendance, your contact getting more junior over time (a signal people miss for years), tone in writing.

💷
Commercial

Procurement appearing early, budget language in routine calls, invoice queries, downgrade questions, the contract read closely for the first time.

🧭
Strategic

Their business, not yours: leadership changes, cost mandates, restructures, M&A, a pivot that makes your use case matter more or less.

Quiet champion+Early invoice query+New CFO
Single signals lie. Convergence tells the truth.

A quiet champion might be on holiday. A quiet champion plus an early invoice query plus a new CFO is a pattern, and patterns across families are how real risk announces itself. Brief your assessments family by family and tell the model to look explicitly for cross-family convergence. It is the difference between an anomaly list and an analysis.

3
Mode A: the single-account deep dive

When an account needs proper attention, a flag fired, a renewal approaches, something feels off, Mode A is the full assessment: everything you have, organised by the four families, into a reasoned analysis. Two instructions make it rigorous. First, reasoning before conclusions (Module 02): you need a logic chain you can audit, not a verdict you must take on faith. Second, and non-negotiable: the disconfirming pass, "What in this data argues against your conclusion? What is the strongest innocent explanation?"

Why the disconfirming pass matters

You usually run an assessment already suspecting the answer, primed to accept whatever confirms it. Forcing the model to argue the other side is the bias correction, built into the prompt where you cannot forget it. Half the time it deflates a false alarm. The other half, the innocent explanations fail and you act with real confidence.

One more discipline: assess trajectory, not snapshot. "Usage is 60% of licence" means nothing on its own. "Usage was 85% two quarters ago" is the actual finding. Date everything you paste.

4
Mode B: the portfolio triage

Mode A does not scale: you cannot deep-dive twenty accounts a week, and the account that kills your year is usually the one you were not deep-diving. Mode B is the answer: the same four families, compressed, across the whole book, every Monday, in thirty minutes. Input: the weekly exports you already have. Output: a ranked watchlist, what changed since last week (the highest-value line in the whole report), and which accounts have earned a Mode A this week.

Mode B
The radar

Compressed scan across the whole portfolio, every Monday. Its job is to never be switched off.

Weekly · 30 minutes · whole book
Mode A
The inspection

Full deep dive on one account the radar flagged. Everything you have, every family, reasoned out.

As needed · deep · one account
The point

A risk process you run brilliantly once a quarter is worth less than a decent one you run every single Monday. Churn signals appear on their schedule, not on yours.

5
Worked example: the account that felt fine
Monday, Mode B flags a convergence

An account you would have called healthy makes the watchlist: reply latency from your main contact has doubled over six weeks (relationship), and finance raised two invoice queries in a month (commercial). Either alone is noise. Two families moving together earns a Mode A.

Tuesday, Mode A with the disconfirming pass

Full data in, families separated, reasoning requested. The risk case: contact disengaging while finance scrutinises spend, classic pre-churn shape, renewal in seven months. Then the disconfirming pass: the strongest innocent explanation is that their publicised finance-system migration explains the invoice queries, and your contact mentioned a department reorganisation in March that could explain the latency.

The output is an action, not a colour

The assessment ends the only way assessments are allowed to end: a named intervention, an owner, a date. "This week: re-engage the contact with a value-led check-in (owner: me, by Thursday) and ask finance directly whether the queries relate to the migration (owner: me, on the same call). If the innocent explanations hold, downgrade. If not, save play, and we have gained a quarter of lead time." Risk that converts to a dashboard colour changes nothing. Risk that converts to an owner and a date changes outcomes.

Practice check · not scored
Read the signals
Four facts about one account: (1) usage flat for two quarters, (2) your champion was promoted and handed you to a junior colleague, (3) a procurement contact asked for the contract terms "for our records", (4) their CEO announced an efficiency programme.

What's the correct reading?

Seniority drift, early procurement interest, and a cost mandate are three different families pointing the same way, which is exactly the convergence the families exist to catch. But convergence justifies investigation, not panic: the promotion could be benign, the contract request routine. Mode A with the disconfirming pass is the proportionate move; both the shrug and the sirens skip the analysis.
Your Mode A concludes: "High risk. Engagement declining, champion disengaged, renewal exposed."

What's the required next line?

Dashboards record risk; they don't reduce it. Reports describe it; meetings discuss it. The non-negotiable standard from this module: every assessment ends in who does what by when. If you can't name the intervention, the analysis isn't finished, however long the document is.

You will leave able to

  • Classify any worry about an account into the four signal families
  • Run Mode A deep assessments and Mode B portfolio triage on a weekly cadence
  • Convert every risk read into a named intervention with an owner and a date

Hands-on exercise

Run Mode B across your real portfolio (anonymised per Module 04). Take the top flagged account into a full Mode A assessment. Compare the result with your gut ranking from before you started, the disagreements are where the learning is.

The human element: a risk score is a hypothesis, not a verdict. The model has never met your champion. Use the assessment to decide where to spend your attention, then go and earn the real answer in conversation.

Includes vault prompts: Churn risk assessment (Mode A) · Portfolio triage (Mode B)

If you remember three things

  1. Churn is a process, not an event. The job is shrinking detection lag.
  2. Single signals lie; convergence across the four families (engagement, relationship, commercial, strategic) tells the truth.
  3. Mode B is the weekly radar, Mode A the inspection, and every assessment ends in an owner and a date.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. You run a churn risk assessment for an account. The AI output says medium risk. You then remember that the product milestone they were waiting for shipped two weeks late and you have not spoken to the champion since. What do you do?

New context changes the output, that is not a failure of the system, that is the system working. A risk assessment built on incomplete input is not a risk assessment; it is speculation with formatting. Add the context, re-run, then act on a complete picture.

2. Your portfolio triage prompt returns all five high-value accounts as green. What is the most important follow-up question?

A triage that only catches acute signals misses churn that builds over quarters, satisfaction drift, silent disengagement, the champion who stopped responding urgently. All green is either genuinely good news or evidence that your system is not measuring the right things. Knowing which it is matters.

3. After completing a churn risk assessment, you ask the model: "Now argue the strongest case that this account is NOT at risk." Which principle does this apply?

You usually run a risk assessment because you are already worried, which means confirmation bias is active before you type the first word. Forcing the model to argue what is wrong with the risk read is the bias correction built into the system. Without this step, you are often just automating your existing fear.

4. You fed the tool everything you know: champion resigned last week, usage down 20% over 90 days, renewal in four months. It still says medium risk. What is the right call?

This is different from re-running with new context. Here the tool already has everything you know, and three serious signals are stacking up at once: a lost champion, falling usage, and a renewal closing in. When the picture is complete and it points one way, the job is to act on your judgment and escalate, not to ask the tool again. Re-running the same facts hoping for a different number is just delay.

Score: 0/4 ·

Tier 2 · Core workflows · Module 08

Expansion and whitespace

Product gap mapping, expansion signals, and business cases the buyer can forward

14 min 85% to pass

The lesson

Same scan,both ways

CSMs are great at spotting danger and oddly blind to opportunity, because risk shouts and opportunity whispers. The reframe that makes this module cheap: expansion signals live in exactly the data you are already scanning for churn. The same Monday scan that protects your book can grow it.

1
The same signals, read the other way

Most CSMs are excellent at spotting danger and oddly blind to opportunity, for an understandable reason: risk shouts and opportunity whispers. A churning account generates escalations and red dashboards. An account quietly ready to grow generates nothing at all unless someone is looking. A new team in the usage logs, a stakeholder from an unfamiliar department joining a call, an initiative announced in their results: the same Monday scan that protects your book can grow it, if you instruct it to read both ways.

The commercial truth

For most mature books, expansion is where a CSM's revenue impact actually lives. Defending the base is the licence to operate. Growing it is the career.

2
The whitespace map

Whitespace is the gap between what they bought and what they could use, but the useful version is not a product checklist. The vault prompt crosses two lists: what you have not sold them against what they have told you hurts, harvested from QBR notes, ticket themes, strategy statements, and stated objectives. Every intersection is a candidate. Most candidates are noise.

The ranking rule is the whole discipline

Order by their problem severity, never by your deal size. The biggest licence uplift attached to a mild inconvenience loses to the modest add-on that removes their team's weekly nightmare, every time. A map ranked by your revenue reads like a sales territory plan. A map ranked by their pain reads like advice. Only one of those gets you invited back.

3
Signals without the ceremony

You do not need a quarterly expansion ceremony. You need three standing questions added to the weekly scan.

👋
Who is new?

Unfamiliar names in usage data or meeting invites mean your product is spreading beyond its original boundary, the single most reliable expansion precursor there is.

📣
What did they just announce?

New initiatives, new markets, new leadership priorities are new problems, and new problems re-rank your whitespace map.

🔧
What are they working around?

Tickets describing manual workarounds, exports into spreadsheets, "is there a way to" questions: each one is a customer describing, in their own words, the gap a product you have not sold them was built to fill.

When a signal matches a whitespace candidate, you have a live opportunity. What you do next is not "tell sales". It is the champion business case.

4
The champion business case

Expansion deals are closed by your champion in a meeting you are not in, armed with whatever you gave them. The deliverable is not a pitch, it is ammunition: a one-page case in their language, against their objectives, with their numbers: the problem as their team feels it, the cost of today's workaround, what changes, the effort honestly stated.

The quality test

Would your champion forward it without editing it? Every sentence that sounds like vendor copy is a sentence they would have to rewrite, and every rewrite is friction between you and the deal. If the case reads like something their own strategy team drafted, it travels upward on its own.

Part 5 of 5
Worked example: from log file to live opportunity
The signal, caught in the Monday scan

Eleven new users appear in the logs, all from a Madrid office that has never used the product. The same week, their interim results mention "accelerating our continental European rollout". Two sources, one story: the whitespace map has your localisation and multi-region module sitting unsold at intersection with "manual region reporting", a pain their ops lead mentioned twice last quarter.

The case, built in twenty minutes

The vault prompt gets everything: the Madrid usage, the results quote, the ops lead's exact words about regional reporting, current workaround costs. Out comes a one-pager titled in their initiative's own name, costed in their hours, with the module as the obvious enabler of a rollout they've already announced. It passes the forward test: nothing in it sells; all of it solves.

The judgement call that stays human

One complication: an open escalation about API performance, eight days old. Strike now or wait? The map can't answer that; the model can't either, honestly. You read it: the escalation is being handled well and visibly, the Madrid team is onboarding this month, and waiting means they entrench the workaround. You brief your champion this week, with the escalation acknowledged in the first paragraph, because pretending it doesn't exist is the one move that would cost you the credibility the whole case rests on. That read, escalation as context to address rather than reason to wait, is yours alone, and Module 10 is about why it always will be.

Practice check · not scored
Rank the whitespace
Three candidates on one account's map: (a) your premium analytics suite, big uplift, they've shown polite interest; (b) a modest workflow add-on that eliminates the manual report their team builds every Friday and has complained about in three separate calls; (c) your newest module, just launched, strategically important to your company this quarter.

Which leads the map?

Their pain ranks the map, full stop. (a) is your deal size talking, (c) is your company's roadmap talking, and the three-at-once proposal is how whitespace maps become ignored attachments. The Friday-report complaint, raised three times without prompting, is the customer writing the business case for you; (b) closing also earns the credibility that makes (a) a real conversation later.
Your business case is ready and strong. This morning, the customer opened a significant escalation about a billing error, theirs to discover, yours to cause.

The timing call?

Timing is the judgement the map cannot make. A billing error you caused, still open, poisons any commercial ask that arrives beside it, but a quarter's delay punishes a live opportunity for a fixable mistake. Fix fast, acknowledge plainly, then proceed: the resolution handled well often strengthens the case. And the model can structure the options, but it isn't in the relationship; the read is yours.

You will leave able to

  • Produce a problem-ranked whitespace map for any account in under 30 minutes
  • Detect expansion signals as a by-product of your existing weekly portfolio pass
  • Draft champion-forwardable business cases in the customer's own language

Hands-on exercise

Build a whitespace map for one real account. Identify the top problem-ranked gap. Draft the one-page case for it. Then apply the test: would your champion forward this unedited? If not, find what's still written for your benefit rather than theirs, and cut it.

The human element: timing an expansion conversation is pure judgement, the model can tell you the case exists, but only you know whether this month's escalation means "not now" or "this is exactly why now". Read the room first; the map will keep.

Includes vault prompts: Whitespace and expansion analysis · Champion enablement pack

If you remember three things

  1. Expansion signals live in the same Monday scan as churn: who's new, what did they announce, what are they working around.
  2. Rank whitespace by their pain severity, never by your deal size. Customers fund pain relief.
  3. The champion case must pass the forward-without-editing test. The timing call is yours alone.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. Your whitespace map identifies a strong product fit for an expansion. What is the most important step before you raise it with the customer?

A well-timed expansion with a healthy account closes. A premature expansion attempt with a strained relationship makes the next renewal harder. The map tells you what is possible; your judgment about the account's current state tells you whether this is the right month. Those are different questions.

2. You ask AI to surface expansion signals from your account notes and call records. What type of signal is it most likely to miss?

AI analyses what is written. Your champion's offhand comment at the end of a call, their visible frustration about a problem that your unreviewed product solves, the thing they said and then added "but don't make a big deal of it", those are yours, and they are often the highest-quality signals in an account.

3. Your champion says "it's not the right time" for expansion. What is the most effective AI-assisted response to this?

"Not the right time" almost always means the conditions are not right yet, not that the case does not exist. A trigger list turns a lost conversation into a monitored opportunity. When the trigger fires, you are already prepared with a case the customer helped define. That is the version that closes.

4. Your AI-generated expansion business case looks excellent. Your champion reviews it and says it looks great. Before presenting it to the executive team, what must you do?

Champion approval means they are comfortable forwarding it, it does not mean it is accurate. One wrong number in an executive business case destroys the entire case, and usually the relationship with it. Every figure needs a source you can name if the CFO asks. "The AI produced it" is not an answer that survives that room.

Score: 0/4 ·

Tier 3 · Mastery · Module 09

Building your AI operating system

From one-off prompts to a personal library and a weekly intelligence cadence

16 min 85% to pass

The lesson

Promptsa system

A prompt saves you forty minutes once. A system saves them every week, feeds its own outputs into its next inputs, and gets sharper as your library matures. Eight modules of techniques now become three components: a library, a cadence, and a set of automation judgements.

1
Prompts are tactics. A system compounds.

Everything before this module made individual tasks faster. This module is about what separates the CSM who "uses AI sometimes" from the one whose entire week runs differently: the first owns some prompts, the second owns an operating system. The difference is compounding. A prompt saves you forty minutes once. A system saves them every week, feeds its own outputs into its next inputs, and gets sharper as your library matures.

2
The library: organised for Tuesday, maintained for March
File by workflow, not by tool

On Tuesday at 14:00 you think "I need call prep", not "I need a Claude prompt". File by workflow: call prep, communication, risk, expansion, internal. The tool note lives inside the entry.

Your library has two failure modes, both organisational. The first is filing by tool: a Claude folder, a Copilot folder. Useless. File by workflow instead. The second failure mode is rot. Models change quarterly, your accounts change monthly, and a prompt that was excellent in January can be quietly mediocre by June without announcing it. The fix costs one line per entry: a last-tested date, plus a note of which model and a link to one example of good output. A library with dates is an asset. A library without them is a folder of expired assumptions wearing the costume of one.

What a library entry looks like
Pre-call brief (example entry)

This is what a maintained library entry looks like in practice. Copy the structure for your own prompts.

Title: Pre-call brief
When to use: Any customer call: QBR, renewal, check-in, EBR
Where it lives: Vault → Call Prep category
Model note: Works well on any approved frontier model
Last tested: [your date here]
Link to good output: [paste a link or file path to your best example output]
Known weaknesses: Misses public intelligence unless layer 2 is added manually
Recent improvements: Added "flag any signal in the four churn families" to constraints (M7)
3
The cadence: the calendar is the system
Schedule beats enthusiasm

Systems live or die on schedule, not on enthusiasm. Two recurring calendar blocks, thirty minutes Monday and ten minutes Friday, do almost all of the compounding work.

The keystone is the Monday portfolio briefing, thirty minutes, before the inbox wins: Mode B triage across the book, expansion questions on (Module 08's three questions), output ranked into "what changed since last Monday". This one habit feeds everything downstream: the accounts flagged become your Mode A list, the briefing context flows into every call prep this week, the expansion signals queue your whitespace work. Highest return of anything in this course, and it costs half an hour you currently spend reacting.

The Monday 30-minute routine
Exactly what to do, in order
8:00
Mode B triage. Open the vault's Portfolio health sweep prompt. Paste your account list with last-contact dates, renewal dates, and any available health scores. Output: ranked list with flags. Time: 10 min.
8:10
Flag the reds and ambers. Any account that moved, or that you have not spoken to in 3+ weeks, goes on the Mode A list for deeper attention this week. Time: 5 min.
8:15
Expansion scan. Run M8's three standing questions across the same account list. New names in usage data? Recent announcements? Workarounds mentioned? Flag any signals. Time: 8 min.
8:23
Set the week's priorities. Three accounts for Mode A attention, two expansion signals to pursue. Write it down. The inbox opens now, you are no longer reacting from nothing. Time: 7 min.

The closing bracket is the Friday self-retro, ten minutes with the vault's coaching prompt: what worked, what you avoided, what one thing changes Monday. It is also when the library gets its maintenance: any prompt that disappointed this week gets its fix folded in while you still remember why. Two blocks, forty minutes total, and your week now has an intelligence system with a heartbeat.

4
Automation: assemble by machine, decide by human

Once the cadence is habit, the temptation is to chain it: exports flowing automatically into triage, flags into briefs, briefs into drafts. Some of that chain is genuinely worth building, and one rule keeps it safe: map the workflow on paper before automating any of it. Every step, every input, every decision point, on one page. If you cannot draw it, you do not understand it, and automating a workflow you do not understand just produces mistakes at machine speed.

The dividing line, always the same

Automate assembly, never judgement. Gathering exports, applying the placeholder pass, running the triage prompt, formatting the watchlist: assembly, automate freely. Deciding which flagged account earns a Mode A, what the intervention is, whether the expansion case goes out this week: judgement, and every judgement point needs a human checkpoint where the chain stops and waits for you. The goal was never a machine that runs your book. It is a machine that sets the table so your judgement is all that is left to add.

Part 5 of 5
Worked example: one week on the operating system
Monday, 08:30, the briefing

Thirty minutes: Mode B flags one risk convergence and one expansion signal (a new department in the usage logs). Two decisions made by 09:00: a Mode A scheduled for the risk account, the whitespace map pulled for the other. Total prompts typed: one, from the library, last tested three weeks ago.

Tuesday to Thursday, the system feeds itself

Tuesday's customer calls are prepped in twelve minutes each, the briefs already half-informed by Monday's context. Wednesday's Mode A ends with an intervention, owner, date. Thursday's QBR uses the narrative-first prompt; the story arc cites usage movements Monday's scan already surfaced. Nothing was started from a blank page all week.

Friday, 16:30, the retro and the maintenance

Ten minutes with the coach prompt. The honest answers: the QBR prompt's output needed too much restructuring (fix folded into the library entry, date updated), and the avoided thing was the awkward pricing conversation with one account (named, scheduled for Monday). The week's ledger: roughly nine hours of assembly work done by the system, every decision still made by you. That ratio, machine-assembled, human-decided, is the whole design.

Practice check · not scored
Design the system
Six months from now, your library has 60 prompts. You run one for a renewal brief and the output is oddly mediocre, though the prompt "always worked before".

What most likely rotted, and what's the fix?

"Always worked before" plus no last-tested date is the rot signature from Part 2: models shifted, your accounts shifted, and the prompt stood still. The benchmark habit exists for exactly this moment. Size isn't the problem (a dated, workflow-filed library scales fine), and tool-hopping rebuilds the same rot somewhere new.
A colleague proposes automating the full Monday chain: exports gathered → placeholders applied → triage run → watchlist formatted → risk emails drafted and sent to flagged accounts' owners with intervention recommendations.

Where must the chain stop and wait?

Assembly automates; judgement doesn't. Everything up to the formatted watchlist is mechanical and safe to chain (the placeholder pass especially: a tested find-and-replace script is MORE reliable than a Friday-afternoon human). But auto-sent intervention recommendations are machine judgements wearing your name, arriving in colleagues' inboxes at machine speed. The checkpoint sits exactly where assembly ends and reading the flags begins.

You will leave able to

  • Build and maintain a personal prompt library with a structure that survives daily use
  • Run a 30-minute Monday portfolio briefing that sets your week's priorities
  • Map and run multi-step workflow chains across your core CS motions

Hands-on exercise

Stand up your library with your five most-used prompts from this course, each adapted to your accounts and tools, each with a worked example and a tested date. Then run your first full Monday briefing on the real portfolio. This exercise is the heart of the whole course: the system you keep.

The human element: systems free attention; they don't replace it. The Monday briefing tells you where to look, the looking, the calling, and the caring remain gloriously manual.

Includes vault prompts: Weekly portfolio briefing · Voice-of-customer synthesis

If you remember three things

  1. File the library by workflow, and date every prompt. Undated libraries rot silently.
  2. The calendar is the system: Monday portfolio briefing, Friday retro. Forty minutes that run the week.
  3. Automate assembly, never judgement, and put a checkpoint exactly where one becomes the other.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. You have 40 saved prompts named "prompt 17" and "renewal final v2." You spend five minutes looking for a specific renewal prompt and cannot find it. What is the actual problem?

A collection you cannot navigate on a busy Tuesday is a graveyard. The fix is not fewer prompts or better tools, it is a naming convention tied to the moment in your workflow when you would reach for it. "CALL-PREP: QBR executive" beats "renewal final v3" every time.

2. You have been running the same Monday portfolio briefing workflow for three months and it takes 25 minutes. You have run it 12 times. What is the single best next step?

Documentation is the step between "my workflow" and "team workflow." A process you can hand off scales to the whole organisation. A Slack post without documentation gives the team the idea but not the system. The documentation is what gives you that reach, and it is also what makes automation safe when you are ready for that step.

3. You want to automate a 40-minute account review process. What is the most important prerequisite before building any automation?

A chain you do not fully understand cannot be debugged when it fails, and it will fail. Understanding every decision point is also what lets you defend the output when someone asks why the risk flag fired on a particular account. Automation before understanding is a liability, not an efficiency gain.

4. A CSM has a prompt library with 30 well-written prompts. They use it twice in three months, even though they work in AI-heavy workflows daily. What is the most likely structural reason?

On a busy Tuesday morning you think "I need call prep", not "I need to browse my prompting section." A library organised around workflow moments, MONDAY-BRIEFING, PRE-CALL, POST-CALL, RISK-ESCALATION, is reachable in seconds. Topic-organised libraries are findable only when you already have time, which is never when you most need them.

Score: 0/4 ·

Tier 3 · Mastery · Module 10

The human element

When not to use AI, how to stay trusted, and turning fluency into career capital

14 min 85% to pass

The lesson

The inversion

When you implement everything in this course, the routine work compresses and what remains is almost entirely the human part: reading rooms, judgement calls, trust. The more fluently you use AI, the more your value concentrates in the things it cannot do, and the more deliberately you must protect them.

1
The inversion

Here is what actually happens when you implement everything in this course: the routine work compresses, and what remains of your job is almost entirely the human part. The reading of rooms. The judgement calls. The trust. Customers were never paying for your typing speed. They were paying for someone accountable who knows them, and that, in an AI-saturated market, just became the scarcest thing a vendor can offer.

2
The do-not-delegate list
The test for the list

Does the value of this communication depend on the customer believing a specific human took the time? If yes, the time is the point, and delegating it deletes the message while keeping the words.

Some work should never be AI-drafted, not because the model would do it badly, but because the authorship is the content. The standing entries:

🙏
The apology after a serious failure

An incident summary can be drafted. The apology that carries it cannot. Customers can survive your product failing. They cannot survive discovering your contrition was generated.

🤝
Live negotiation

Rehearse beforehand relentlessly, but the room is yours alone. Reading a pause, sensing when to hold silence, deciding in the moment to give ground: improvising from a machine mid-conversation outsources the one skill the conversation exists to test.

💬
Anything about a person's situation

A contact's redundancy, a champion's illness, a congratulations that matters. Three human sentences beat three perfect paragraphs.

🔧
Relationship repair

When trust itself is the broken thing, the labour of the message is the repair.

Notice what is not on the list: almost everything else. The list works precisely because it is short, a few protected categories fiercely held, while the other 90% of your output gets the full leverage of the previous nine modules.

3
The question, and the answer you have earned

Sooner or later a customer asks it, sometimes curious, sometimes pointed: "Did you write this?" The two losing moves are denial (one metadata accident from being a trust crisis) and apology (which concedes something improper happened).

The winning answer this whole course has been building

"I use AI the way I would use a junior analyst: it assembles drafts and does the legwork, and every judgement, every fact, and every word that reaches you is mine, because I edited it, verified it, and chose to send it."

That answer usually improves the relationship, because the real question was "am I still dealing with someone accountable?", and you have just answered yes with evidence. You are entitled to it exactly as long as the workflow behind it is true, which is why Module 04's output discipline was never a compliance chapter. It was the foundation of this sentence.

4
Turning fluency into career capital

You now run a measurably different operation from most of your peers. Two moves convert that into career value, and secrecy is not one of them.

📊
Quantify your story

The narrative leadership remembers is specific: "I recovered roughly six hours a week, and reinvested them in face time with my top accounts. Here are the two saves and the expansion that came out of it." Recovered hours are the input. The saves and growth are the story.

🎓
Teach it

Run the lunch-and-learn, share the library, walk a colleague through their first Mode B. Giving the methods away beats hoarding them: the hoarder is a CSM with a trick, the teacher becomes the person the organisation associates with the whole capability.

Part 5 of 5
Worked example: the question, live
The setup

A QBR ends well. As laptops close, the customer's COO, sharp, friendly, slightly testing, says: "These briefs you send are impressively thorough. AI, presumably?" Two colleagues look up. The next ten seconds are the module.

The losing versions

"No, all me." A denial you now have to maintain forever, one forwarded email with a paste artefact from being a credibility event. Or the nervous over-apology: "Yes, sorry, I should have mentioned...", conceding fault where none exists, and teaching the room that your tools are something to be embarrassed about.

The earned version

"Yes, for the assembly. I use it like a junior analyst: it pulls the data together, and then I do what you're actually paying for, the judgement about what matters and what we should do about it. The recommendation on slide nine, that came from me reading your reorganisation, not from a model." The COO nods, and the conversation moves on, except it doesn't quite: you've just positioned yourself as the vendor contact who is ahead of the technology rather than hiding behind it. Six months later, when their team is debating its own AI policy, guess who gets asked for a coffee.

That's the course. Ten modules, one thesis, proven in the last place it gets tested: AI for the assembly, you for the judgement, and the human element not merely intact but more visible than it ever was when it was buried under forty minutes of call prep. Finish the knowledge checks, claim the certificate, and go be the CSM the next five years belong to.

Practice check · not scored
Hold the line
Friday, 17:20. You learn your favourite champion, eight years at the account, has been made redundant, announced internally an hour ago. You want to send something tonight. The vault has a prompt that would draft a warm, well-structured note in thirty seconds.

The right move?

A person's redundancy is squarely on the list: the note's entire value is that a human who cares spent the time, and "lightly edited" doesn't change who did the caring. Even the polish pass fails here, because what makes the message land is precisely its imperfect, unmistakably-you texture. Three honest sentences tonight beat anything generated, ever. The other 90% of your week gets the machine; this is the 10% that never does.
Performance review season. You've run the operating system for six months: roughly six hours a week recovered, one at-risk account saved, one expansion sourced from a Monday scan signal.

Which version of the story builds career capital?

Outcomes, mechanism, multiplication, in that order. "Proficient with AI" is a skills-matrix checkbox; tooling detail is a hobby presentation; silent results get attributed to luck or market. The winning story connects recovered hours to commercial outcomes and then offers to scale it through the team, which converts a personal edge into organisational value with your name on it.

You will leave able to

  • Apply the do-not-delegate list without exception, and explain why it exists
  • Answer "did AI write this?" in a way that strengthens the relationship
  • Present a quantified AI ROI story to leadership and convert it into career capital

Hands-on exercise

Write your one-page AI ROI story: hours recovered per week, what you reinvested them in, one concrete outcome it produced (a save, an expansion, a faster cycle). Then book the meeting where someone senior hears it. The course ends when that meeting is in the diary.

The human element: this whole module is the human element. The promise of this course was never "do less work", it was "do more of the work only you can do". If your customers trust you more after this course than before it, you've passed.

Includes vault prompts: Objection and negotiation rehearsal

If you remember three things

  1. The inversion: the more you automate, the more your value concentrates in what AI cannot do.
  2. Keep the do-not-delegate list short and fierce: when the authorship is the content, the labour is the message.
  3. Career capital comes from outcomes, mechanism, multiplication: quantify your story and teach the team.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

4 questions · 85% to pass

1. Which of these should never be delegated to AI, even with the strongest available prompt?

Engagement metrics look healthy right up until they do not. The customer who stopped escalating because they gave up, not because things improved, is a pattern that requires reading what is absent: the calls that did not happen, the energy that left the room, the questions they used to ask that went quiet. That read is yours, not the model's.

2. A customer says: "20% reduction or we're not signing." What is the right role for AI in this moment?

Negotiation is presence, reading, and live judgment. AI helps you walk in fully prepared, every scenario worked through, every objection already heard and answered in practice. In the room, it is you. The CSM who has rehearsed every version of this conversation does not need a script; they need the mental freedom that comes from preparation.

3. A customer says: "You clearly understand our situation better than any vendor we work with." What is the highest-trust response to this moment?

The document showed you prepared. Walking through your thinking live proves you understood, and that you can do it without a script. That distinction is the difference between AI-assisted and genuinely skilled, and customers who work with CSMs long enough learn to tell the difference.

4. A junior CSM on your team is producing strong renewals but cannot handle escalations without running AI prep first, and struggles visibly when conversations go off-script. As their leader, the right response is:

Dependency is only visible in failure modes, and by then the relationship damage is done. The deliberate practice design: work through the escalation without AI, commit to your approach, then use AI in debrief to see what you missed and why. This builds underlying judgment rather than managing dependence, and produces a CSM who is genuinely skilled rather than just AI-enabled.

Score: 0/4 ·

Why this exists

You don’t have a capacity problem.
You have an AI problem.

45 minutes of meeting prep. A QBR deck built at 9pm. A churn that "came out of nowhere" but had been sitting in your usage or support ticket data for a quarter. That isn’t a capacity problem. It’s assembly work AI should be doing for you, and learning to hand it over properly is exactly what this course teaches.

  • The 45-minute meeting prep takes a fraction of the time
  • Churn stops coming out of nowhere
  • No more late-night QBR prep and it reads like you wrote every word
  • When a stakeholder goes quiet, you know what to do next
What this looks like in practice

From reactive to in control
in one working week.

Before
  • Monday: 18 accounts open in the CRM. Couldn’t hold them all in my head. Prepped the two calls in the diary, deferred the rest.
  • Tuesday: customer asked about an open feature request mid-call. Hadn’t read the ticket thread. Promised to come back.
  • Wednesday: manager asked “any risks I should know about?”. Took ninety minutes to give a confident answer.
  • Friday: spotted a sentiment shift in a Thursday email thread I should have caught the day before. Apologised, replied late.
Reactive · always one step behind
After
  • Monday: 30-minute portfolio sweep flagged the three accounts to prioritise. Briefs ready for every customer-facing call before the diary opened.
  • Tuesday: walked in with the open ticket already in the brief. Answered the feature question in the moment, in context.
  • Wednesday: manager’s question answered in five minutes. Two risks named with the evidence behind each.
  • Friday: end-of-week inbox sweep caught the sentiment shift the day it landed. Same-day reply, follow-up booked.
In control · earlier on every signal

See the prompts that made this happen →

Watch it happen

Same task. Same AI. Same data.
The only variable is the prompt.

How most people promptGeneric
The AI-Powered CSM MethodSpecific

One of these walks into the renewal call ahead. The technique is in Module 02.

The Direct/Present Method

Two disciplines.
One shift. A whole different week.

Every CSM role, at every scale, does two fundamentally different kinds of work. Work that can be assembled, and work that must be present. The Method is how you handle both, and manage the transition between them.

DISCIPLINE 1
Direct
work that can be assembled

Briefs, decks, status notes, risk write-ups, first drafts. You direct AI to produce; you verify what comes back. Four practices in order: Diagnose, Brief, Layer, Verify.

THE SHIFT
The transition where it usually breaks
DISCIPLINE 2
Present
work that must be present

Live customer moments. Judgement calls. The moment a hard question lands. AI can assist (transcribing, retrieving) but never substitutes for your attention or your account of yourself.

Direct the assembled work. Be present for the moments that require it.
Never let one mode carry into the other.

Read the full Method →

Stuck on an account right now?

The Method is how to think. Situation Read is the thinking, applied. Tell it the situation you are actually in, a customer gone quiet, usage dropping, a renewal that looks fine but feels off, and get the read a working CSM would give you: what is really going on, the one move this week, and what to watch. It adapts to your level, and nothing you tap ever leaves your browser.

Get the read on your situation →
The build

Build your own AI team.

I run a team of specialist agents across a live enterprise portfolio. They watch my accounts, inbox, calendar and the wider world, and surface the handful of things that matter each day. They read everything and touch nothing. Every decision stays mine.

This is the full blueprint: the five roles, the hard lines, and how to build your own version whether you have a connected setup or just a single chat. Built on one principle: human as the loop, not human in the loop.

See how it is built →
Certification

Earned, not attended.

Most course certificates prove you watched videos. These prove you changed how you work. Three courses, three separate certificates, same standard. Made to live on your LinkedIn profile, and to hold up if your manager asks how you earned it.

1

Work through a course

Lessons, drills, and worked examples, with exercises that put each module to work on your live portfolio. Self-paced and built to fit around the day job.

2

Pass every knowledge check

Each module ends in a short scenario quiz where wrong answers are explained, not just marked. A module only counts as complete once its check is passed, so the certificate means what it says.

3

Generate your certificate

Finish every module in the course and the certificate unlocks. Enter your name and download it instantly, dated and stamped with a unique ID. Made to go straight on your LinkedIn profile.

Sample

A live render of the Foundation certificate. Advanced and Leaders follow the same standard. Yours downloads in full resolution as a landscape image and a square version made for LinkedIn.

From the founder

Why I built this

Gary Giacalone, founder of The AI-Powered CSM

I'm Gary Giacalone, a Customer Success Manager just like you. I created The AI-Powered CSM because I saw a gap between the AI training available and what Customer Success professionals actually need.

Most AI courses are either too generic or too technical. What was missing was practical guidance for the real work CSMs do every day: preparing for renewals, building QBRs, analysing account health, writing customer communications and finding more time to be strategic.

The same concerns kept coming up

  • "I know I should be using AI more, but I don't know where to start."
  • "How do I know if I can trust the output?"
  • "Which tools are approved and what data can I actually share?"

Those are valid concerns. The goal of this course isn't to replace the skills that make great Customer Success professionals successful. It's to help you strengthen them.

Gary Giacalone · Founder, The AI-Powered CSM

Read the monthly newsletter →
My aim is simple

Ensure every Customer Success professional becomes confident with AI, without losing the human element that makes them good at their job.

The CSMs who learn this now will spend the next 5 years ahead.

Start The AI-Powered CSM, free
The AI-Powered CSM

Built for CSMs by a CSM

© 2026 The AI-Powered CSM · Free for the Customer Success profession · Your progress stays in your browser, nothing is tracked Verify a certificate →
First of its kind · instant · nothing stored

The Prompt Grader

← Back to the Prompt Vault

Paste a prompt you actually use. Pick your AI tool, and it gets scored live against the RCCO framework from Module 02, plus how well it fits the way your chosen tool likes to be briefed, with the exact fixes.

Scoring low? Module 02 teaches all four components, and every prompt in the vault has them built in for all four tools.

Free · 2 minutes

Your AI Fluency Score

Eight real CSM scenarios. No sign-up, nothing stored. You get a score out of 100, your profile across four skills, and exactly where to start.

AI Governance

Not a policy doc.
A working system.

Six rules, one fast check, a data matrix, and four real scenarios. Everything a CSM needs to use AI professionally and keep customer trust intact.

6core rules
4safety checks
4scenarios
1data matrix
Quick check

Can I paste this?

Run before anything customer-related goes near a model. Takes 20 seconds.

Interactive check

Can I paste this?

Answer honestly about whatever is on your clipboard right now.

Reference

What can go where

Not all tools carry the same protections. Not all data carries the same risk.

Data type
Enterprise / approved
Business subscription
Free / personal
Customer names & contactsEmails, phone, LinkedIn, org chart
Anonymise first
Anonymise first
Never
Account health & metricsARR, health score, usage data
Safe
Check DPA
Never
Call transcriptsRecorded meetings, Teams/Zoom exports
Anonymise first
Anonymise first
Never
Pricing & strategyDeal economics, roadmap, NDA material
Anonymise + clear
Never
Never
Anonymised / placeholder dataCUSTOMER_A, CONTACT_1_IT, $XXXK ARR
Safe
Safe
Fine
Your process & templatesPrompt libraries, workflow docs, frameworks
Safe
Safe
Fine

Rules assume standard enterprise AI terms. Always verify against your organisation's DPA and customer contracts.

Practice

Would you catch this?

Four real situations every CSM faces. Tap what you'd do.

The rules

Six rules that don't change

The tools change every six months. These don't.

For managers

Team policy template

Copy, fill in the blanks, send before your team uses AI with customer data.

AI usage policy, CS team
Approved tools: [LIST YOUR APPROVED TOOLS, e.g. Claude Enterprise, Copilot M365] Before using AI with customer data: 1. Approved tools only, not free tiers or personal accounts 2. Replace names with placeholders: CUSTOMER_A, CONTACT_1_IT 3. Pricing strategy, legal positions and NDA material: approved tool only, anonymised and cleared first 4. Verify every figure before it reaches a customer 5. If unsure, anonymise, the analysis quality is identical Questions? Ask [YOUR NAME].
Escalation

Three questions for IT or legal

Asking these marks you as the person who gets it.

01
Which AI tools are approved for customer data, and at which subscription tier?
This is your first question, before your next prompt, if you don't already know the answer.
02
Do our customer contracts or DPAs restrict processing customer data with AI sub-processors?
Enterprise contracts often have clauses that go further than the tool's standard terms.
03
Is there a review path for new AI tools, and who owns it?
You want to be the person who routes requests, not the one who went rogue.

The full discipline is in Module 04: Data hygiene and AI safety.

For CS Leaders

Your team is using AI.
The question is how well.

A common framework, shared vocabulary, and 77 live prompts. Here is how to deploy it as a leader.

3-5hper CSM
10modules
77live prompts
Freeto access
The business case

What your team gets

Concrete outcomes, not AI awareness.

Pre-call prep
A fraction of the prep time

Account briefs that used to mean trawling CRM, email and tickets by hand now come from a single prompt, the pre-call intelligence brief, in a fraction of the time.

Portfolio scan
Signals, not hunting

The Monday portfolio briefing and dual-mode churn triage replace manual health-score review, so CSMs spend their time acting on signals rather than digging for them.

QBR quality
Story-first decks

The course teaches narrative-before-slides. QBRs that used to fill from a template now open with a customer-specific story arc.

Rollout

Four steps to deploy this

No LMS. No licences. Copy the message below and go.

01
Have everyone take the AI Fluency Score first
2 minutes. Baselines where each person starts. Reference it in your next 1:1. Link: ai-powered-csm.com
02
Assign Modules 01-04 as the foundation sprint
~90 minutes. Mental model, prompting framework, tool selection, data safety. Everything after builds on them.
03
Run a team prompt review in your next meeting
Each CSM brings one vault prompt they ran on a real account. Twenty minutes of shared, real-account learning beats reading alone.
04
Re-score the team after 30 days
Fluency Score again. The deltas show who applied it and where gaps remain. Use it to direct coaching, not rank people.
Send this

Team launch message

Slack / email template
Hey team, I want us to build a consistent approach to AI. Not rules, a shared way of working. I am asking everyone to complete The AI-Powered CSM course over the next two weeks: ai-powered-csm.com Start with the AI Fluency Score (2 minutes), then work through Modules 01 to 04. In our next team meeting, bring one prompt you ran on a real account and what came back. No pressure on pace. But do the score first.
Governance

Protect your team first

Share the AI Governance section before your team uses AI with customer data. Data risk matrix, interactive check, and a policy template they can keep open in a tab.

Open AI Governance →
One-click deployment assets

Everything your team needs to start this week.

Copy and adapt these assets. Brief your team, set the calendar, make the business case, and send the Slack message in one sitting.

📋

Team Briefing Doc

A ready-to-send document explaining what AI-Powered CSM is, why you are rolling it out, and what you expect from each team member.

Copy briefing →
📅

4-Week Rollout Calendar

Week-by-week schedule: which modules each week, when the group debrief happens, and how you track adoption.

Copy calendar →
📊

Business Case Template

Two-page business case for your CFO or VP. ROI framing, time-saving estimates, and risk of inaction.

Copy business case →
💬

Pre-Written Slack Message

Drop this into your team channel today. Warm, direct, and sets the right tone without overselling it.

Copy message →
The Prompt Vault

The right prompt,
for the right moment.

Advanced · build with AI

Become the most AI-capable person in your CS team.

Twelve modules that take you from using AI to building with it: your own digital twin, custom assistants, a working agent, and the authority that comes with them. Harder than the foundational course, on purpose.

Build
Judgement
0/12

Modules complete

No account, no tracking, nothing leaves your machine. Your progress lives in this browser, so stick to the same browser to keep your place. The knowledge checks pass at 85 percent.

Part 1 · Build yourself

From one-off chats to a real personal system

M01–M05 · prompt library, your AI twin, instructions that hold

MODULE 01

Stop using AI one chat at a time

The shift from one-off chats to systems you reuse, and how to spot which is which

MODULE 02

Turn your messy prompt folder into a real library

Naming, versioning, and testing your prompts so they stay sharp instead of quietly rotting

MODULE 03

Build a version of AI that works like you, part one

Capturing how you actually work, your standards, your voice, your judgement, in writing

MODULE 04

Build a version of AI that works like you, part two

Giving your twin the knowledge it needs, in Claude, ChatGPT, Gemini, or Copilot, safely

MODULE 05

Write instructions that actually hold up

Why assistants drift, and how to write rules that stay steady under pressure

Part 2 · Build systems

From assistant to agent, safely

M06–M09 · agents, judgement, chaining, all with guardrails

MODULE 06

From an assistant that answers to an agent that acts

What an agent really is, where the builders are today, and guardrails before power

MODULE 07

Build a real CS agent from start to finish

One genuinely useful agent, from idea to tested tool you could defend to your VP

MODULE 08

Judge AI output like an expert

Catching confident-but-wrong output, and the honest truth about when to trust the tool over your gut

MODULE 09

Join the pieces into something that runs itself

Chaining your builds into a reliable routine, with checks so one failure does not cascade

Part 3 · Lead

Turn your work into authority

M10–M12 · be the AI person, defend what you built, teach it

MODULE 10

Become the AI person for your whole company

Building a shared CS knowledge base, and owning the accuracy that makes you the authority

MODULE 11

Keep what you built working, and defend it

Testing, catching silent decay, and defending any output when someone pushes back

MODULE 12

Turn what you built into real standing

Sharing, teaching, and positioning so your work becomes a reputation, not a private trick

Part 1 · Build yourself · Module 01

Stop using AI one chat at a time

The shift from one-off chats to systems you reuse, and how to spot which is which

12 min 85% to pass

The lesson

1shift

Most people open a chat, ask for something, get it, and close the tab. The next day they start from scratch. This whole course rests on one shift: stop using AI one chat at a time, and start building things that work again and again without you rebuilding them.

1
Why one-off chats keep you a beginner

Here is the quiet ceiling most people hit. They get good at writing a prompt, they get a good answer, and then the answer disappears into a closed tab. The next time the same job comes round, they start again from a blank box. They are fast at a task that should not exist any more, because they keep rebuilding it.

The shift that separates the advanced operator from the rest is not better prompts. It is realising that most of your work is the same handful of jobs, repeated, and that each of those jobs can be built once and reused. You stop being someone who uses AI and start being someone who builds with it. That is the entire difference, and everything in this course is a version of it.

The point

A good prompt saves you forty minutes once. A built thing saves you forty minutes every single time the job comes round, and it gets sharper each time you touch it. The first is a trick. The second is leverage.

2
Saved prompt, assistant, or agent

There are three things you can build, and most confusion comes from not knowing which one a job actually needs. They are not the same, and reaching for the wrong one wastes effort or, worse, hands real decisions to something that should never have had them.

📝
A saved prompt

A reusable instruction you keep and paste when you need it. Best for a self-contained job you do often: the pre-call brief, the QBR narrative, the ticket sweep. You are still in the chat, still in control, just not rewriting the instruction each time.

🤖
A custom assistant

A saved prompt plus knowledge plus a fixed character: a Claude Project, a Custom GPT, a Gemini Gem, a Copilot agent. It already knows your context and works in your voice, so you brief it far less. Best for a role you play often, like your own drafting twin.

⚙️
An agent

Something that works towards a goal across several steps and takes real actions, not just answers. Best for a repeatable process, like watching your portfolio and drafting an alert. The most powerful and the most dangerous, because it acts.

The test

Does the job need you in the loop every time (saved prompt), a standing helper that knows you (assistant), or a process that runs and acts on its own (agent)? Match the build to the job. A renewal email is a saved prompt. Your drafting voice is an assistant. Monday portfolio triage that flags and drafts is an agent.

3
Spotting what is worth building first

Not everything is worth building. The honest filter is two questions: how often do you do this job, and how much does it cost you each time? High frequency and high cost is where you build first. Low frequency, even if painful, often is not worth the setup. A beautiful agent for a job you do twice a year is a hobby, not leverage.

There is also a trap on the other side: building something complex for a job a saved prompt would handle. Complexity has a cost too, things that act can fail in ways things that answer cannot. The advanced operator is not the one who builds the most. It is the one who builds the right thing for each job, and leaves the rest as simple chats.

What you walk away with

A simple map of everything you do over and over in a week, each item scored by frequency and cost, so you know exactly what to build first and what to leave alone. That map is the plan for the rest of this course.

The job that should be a saved prompt

You write a post-call follow-up email after most calls. The shape is always the same, but the content is specific to each call, so it needs you every time. This is a saved prompt: a reusable instruction with slots for the call details. Building an assistant would be overkill, and an agent should never send a customer email unsupervised.

The job that should be an assistant

You draft a lot of writing and you have a distinct voice your customers know. Re-teaching the model your voice every time is the waste. This is an assistant: a Project or Custom GPT loaded with samples of your writing and your standards, so it drafts in your voice from the start. You still edit, but you start from a much higher baseline.

The job that should be an agent

Every Monday you pull exports, scan for risk and expansion signals, and rank your book. The steps are the same each week and the gathering is mechanical. This is an agent: it assembles the data, runs the analysis, and drafts the watchlist, then stops and waits for you to make the judgement calls. Assembly by machine, decisions by you.

Practice check · not scored
Try the judgement call
A CSM says: "I write a renewal risk summary for each of my 30 accounts every quarter. It is the same structure each time but the facts differ per account, and I read every one myself before it goes to my manager."

What should they build?

The job is frequent and structured, so it is worth building, which rules out doing nothing. But they read every one before it goes up, so a human is required in the loop every time, which rules out an autonomous agent. A saved prompt with a fixed structure, run per account, is the right weight. The assistant-with-all-data option sounds clever but mixes 30 accounts' context together, which hurts accuracy and raises data questions.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. You do a particular AI-assisted job about twice a year. It is fiddly and mildly annoying each time. A colleague suggests you build a custom assistant for it. What is the best call?

Frequency is half the filter and this job fails it. Twice a year means the time you spend building and maintaining the thing will likely exceed the time it ever saves. The honest advanced move is restraint: a normal chat is the right tool for rare jobs. Building for the sake of it is a cost, not a win.

2. Which of these jobs is the strongest candidate to build as an agent rather than a saved prompt or assistant?

An agent suits a repeatable, multi-step process where the steps are stable and most of the work is assembly, which the weekly routine is exactly. The renewal email and the apology are human-in-the-loop or do-not-delegate work. Drafting in your voice across varied tasks is the textbook case for an assistant, not an agent, because there is no fixed multi-step process to automate.

3. What is the real reason building everything as a complex agent is a mistake?

The key insight of this module is that the three build types carry different risk, not just different effort. An agent acts, so a mistake becomes an action taken in the world, not just a bad paragraph you can ignore. That is why matching the build to the job matters: you do not hand action to something that only needed to answer. Cost and speed are minor next to that.

Score: 0/3 ·

You will leave able to

  • See your work as a set of repeatable jobs, not one-off chats
  • Tell a saved prompt, an assistant, and an agent apart, and pick the right one
  • Score your jobs by frequency and cost so you know what to build first, and what to leave alone

Hands-on exercise

List every AI-assisted job you did last week. Mark each one as rare or frequent, and cheap or costly per time. Circle the frequent-and-costly ones: those are your build list for this course. Then label each as saved prompt, assistant, or agent.

The human element: Building is leverage, not a goal in itself. The most fluent operators build deliberately and leave plenty as simple chats. Reaching for the heaviest tool every time is a beginner's tell, not an expert's.

If you remember three things

  1. Stop using AI one chat at a time. Most of your work is a few jobs repeated, and each can be built once and reused.
  2. Three things to build: a saved prompt (you stay in the loop), an assistant (a standing helper that knows you), an agent (a process that acts).
  3. Build the frequent, costly jobs first. Match the build to the job, and never hand action to something that only needed to answer.
Part 1 · Build yourself · Module 02

Turn your messy prompt folder into a real library

Naming, versioning, and testing your prompts so they stay sharp instead of quietly rotting

12 min 85% to pass

The lesson

40prompts, none findable

Most people end up with forty prompts called things like "renewal final v2" and can never find the one they want when they are busy. A real library fixes that: organised around the moments in your week, versioned, and tested so you notice when a prompt that used to work has quietly stopped working.

1
Why a folder of prompts is not a library

Saving prompts feels like progress, and it is a start. But a folder of forty prompts with names only you half-remember is not an asset, it is a junk drawer. On a busy Tuesday at two o'clock you do not want to scroll through "renewal v2 FINAL actual" and "good churn one". You want to reach for the right tool in five seconds and trust it still works.

A library has three things a folder does not: a naming system you can navigate without thinking, versions so you can improve a prompt without losing the old one, and tests so you know each prompt still does its job. Miss any of these and the library slowly rots into a folder again.

The point

A prompt library is not a place you store prompts. It is a system that keeps your best prompts findable, improvable, and trustworthy over time. The storing is the easy part. The keeping-good is the work.

2
Name and file for the moment you will reach for it

The single most common filing mistake is organising by tool: a Claude folder, a ChatGPT folder, a Copilot folder. It feels tidy and it is useless, because on Tuesday you do not think "I need a Claude prompt", you think "I need to prep this call". File by the job, by the moment in your week: call prep, communication, risk, expansion, internal reporting. The note about which tool works best lives inside the entry, not in the folder name.

Each entry needs more than the prompt text. A good entry carries its title, when to use it, which tool it works best on, the date you last tested it, a link to one example of good output, and any known weaknesses. That sounds like a lot, but it is the difference between a prompt you trust and a prompt you have to re-check every time you use it.

The filing rule

File by workflow, never by tool. The question you ask when reaching for a prompt is about the job in front of you, not the software. The tool note belongs inside the entry.

3
Version and test, because prompts rot

This is the part almost nobody does, and it is what separates a real library from a hopeful one. Prompts rot for two reasons. Your accounts and your job change month to month, so last quarter's perfect prompt slowly stops matching reality. And the models themselves change: a model update can quietly change how a prompt behaves, sometimes for the better, sometimes worse, and it will not announce it.

The fix is cheap. Keep a version when you make a real change, so you can roll back if the new one is worse. And give every important prompt a test: a saved real (anonymised) input and a note of what good output looks like. When a model updates, or every quarter, re-run your handful of core prompts against their test and check the output still holds. Twenty minutes, and you catch the quiet rot before it embarrasses you in front of a customer.

The discipline

Every important prompt gets a last-tested date and a saved example of good output. Anything untested in a quarter, or right after a model update, gets re-run against its test. A library with dates is an asset. A library without them is a folder of expired assumptions.

What happened

Your churn-risk prompt was excellent in January: clear reasoning, caught the convergence, ended in an action. In June you run it and the output is vaguer, more hedged, and skips the disconfirming step you relied on. Nothing in your prompt changed. What broke?

How the library finds the cause

Because you kept a saved test input and an example of January's good output, you can run the exact same input again and compare side by side. You try it on your current model and on the version you used in January. If the old model still produces the good output and the new one does not, the model update changed the behaviour. If both now underperform, something in your accounts or your expectations has drifted, and the prompt itself needs an update.

The fix

Once you know it is the model, the fix is usually a small prompt change: re-state the disconfirming step as an explicit, non-optional instruction rather than relying on the model to remember it. You save that as a new version, note the date and model, and your test now guards against the same drift next time. Without the saved test and example, you would never have known it broke until a bad assessment reached your manager.

Practice check · not scored
Try the judgement call
You organised your prompt library into folders named "Claude", "ChatGPT", and "Copilot". A colleague says it will not work well in practice.

Why are they right?

Filing by tool fails because it does not match how you search in the moment. You reach for a prompt by the job you are doing, so the library should be organised by workflow: call prep, risk, expansion, communication. Which tool a prompt runs best on is real information, but it belongs inside the entry, not as the top-level structure. Prompts are also not incompatible across tools, so that reasoning is wrong even though the conclusion to avoid tool-folders is right.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. A prompt that gave great output in March now gives weaker output in June, and you have not changed a word of it. What is the most useful first step to find the cause?

The whole point of keeping a saved test input and example of good output is to make this diagnosable instead of guessing. Running the same input on both models isolates the variable: if the old model still produces the good result and the new one does not, it is a model change, if both fail, it is drift in your accounts or expectations. Rewriting or switching tools blindly skips the diagnosis and might fix the wrong thing.

2. What is the single most valuable thing to record in each library entry, the thing that turns a folder into a real library?

The defining feature of a library over a folder is that it stays trustworthy over time, and that requires knowing whether each prompt still works. A last-tested date plus a saved example of good output is what lets you spot rot and re-verify. The creation date, usage count, and provenance are nice-to-haves that tell you nothing about whether the prompt still does its job today.

3. Why is keeping versions of a prompt (rather than just overwriting it when you improve it) worth the small effort?

Versioning exists because not every change is an improvement, the same reason developers keep version history. When you adjust a prompt and the output gets worse, you want to return to the known-good version instantly rather than trying to reconstruct it from memory. There is no legal requirement, more versions do not inherently improve anything, and the model does not read your version history.

Score: 0/3 ·

You will leave able to

  • Set up a library you can navigate on a busy Tuesday, organised by workflow not by tool
  • Keep versions and tests so you know each prompt still does its job
  • Spot when a model update has quietly broken a prompt, and fix it deliberately

Hands-on exercise

Take your five most-used prompts. Give each one a proper entry: title, when to use it, best tool, today's date as last-tested, and a saved example of good output from a real anonymised input. You now have the start of a real library and a test you can re-run next quarter.

The human element: A library is a maintenance habit, not a one-time build. The CSM with ten well-tested prompts beats the one with eighty untested ones every time, because they can actually trust what they reach for.

If you remember three things

  1. A folder stores prompts. A library keeps them findable, improvable, and trustworthy over time.
  2. File by workflow, never by tool. The tool note lives inside the entry.
  3. Prompts rot, from your accounts changing and from model updates. A last-tested date and a saved example of good output are how you catch it.
Part 1 · Build yourself · Module 03

Build a version of AI that works like you, part one

Capturing how you actually work, your standards, your voice, your judgement, in writing

12 min 85% to pass

The lesson

1of you

A digital twin sounds fancy, but it just means an assistant that thinks and works the way you do: your standards, your voice, the things you always check, the way you make calls. The first step is putting your own way of working into words, which most people have never actually done.

1
What a digital twin really is

Strip away the science-fiction word and a digital twin is simple: a custom assistant set up so its output sounds and reasons like you, not like a generic helpful stranger. When you ask a fresh chat to draft a renewal note, you get the average of every renewal note ever written. When you ask your twin, you get something close to what you would have written, because it carries your standards, your voice, and the checks you always run.

This module is only the first half: writing down how you work. The second half, in the next module, is giving it your knowledge. We split them because the writing-down is the part people skip, and it is the part that matters most. An assistant loaded with documents but no sense of how you operate is still a stranger with a filing cabinet.

The point

Your twin is only as good as your description of how you work. The model already knows how to write. What it does not know is how you write, what you would never say, and what you always check before anything goes out. That description is the whole build.

2
Putting your method into words

Here is the uncomfortable part: most experienced CSMs have never actually articulated how they work. They just do it. So the real task of this module is self-knowledge, getting the thing in your head onto the page. There are four things worth capturing, and being specific matters far more than being thorough.

🎯
Your role and standards

Who you are and what good looks like to you. Not "a helpful CSM" but the specific bar: "I never send a risk read without a named action and owner. I always check trajectory, not just the snapshot." The standards are what make it behave like an expert, not an intern.

🗣️
Your voice

How you actually sound. Direct or warm? Short sentences or flowing? Do you open with the point or build to it? The fastest way to capture this is to paste examples of your real writing and say "write like this", rather than trying to describe it.

The checks you always run

The things you do without thinking before work goes out: verify every figure, hunt for promises you did not authorise, read it as the customer's legal team would. These become the assistant's built-in habits.

🚫
What you would never do

Your hard lines. Never invent a statistic. Never use the word "partnership" with this kind of account. Never send an apology you did not write yourself. The things it must refuse are as defining as the things it should do.

Specific beats thorough

A short description full of your actual specifics beats a long generic one every time. "Write like these three real emails of mine" teaches more than a paragraph of adjectives. The model fills gaps with the average, so every gap you leave generic comes back generic.

3
Why the generic-assistant trap is the main risk

The most common failure when people build their first twin is that it sounds like everyone else's. They write instructions like "you are a professional, helpful customer success manager who communicates clearly" and wonder why the output is bland. That instruction describes eighty per cent of CSMs. It gives the model nothing specific to anchor to, so it produces the average.

The fix is to write the things that are true of you and not of most people. The phrase you overuse and want kept. The structure you always follow. The opinion you hold that others do not. Specificity is the entire game: the parts of your description that are unique to you are the only parts that make the twin sound like you rather than like a competent stranger.

What you walk away with

Your own written operating instructions: your role, your standards, your voice with real examples, your standing checks, and your hard lines. This is the brief you will load into an actual assistant in the next module.

The generic version

"You are an experienced, professional Customer Success Manager. You write clear, friendly, professional communications. You are helpful and thorough. Always maintain a positive and professional tone." Every sentence here is true of almost every CSM alive. The model has nothing to grip, so it produces polished, generic output that sounds like a vendor template.

The sharp version

"I open with the point, never with pleasantries. I write short. I never use the words 'partnership', 'leverage', or 'reach out'. Before any risk note goes out I name a specific action, an owner, and a date, or it is not finished. I would rather be slightly blunt than vague. Match the voice of the three example emails below." Every line here is a real, specific choice that is not true of most people.

The difference in output

Fed the same task, the generic twin produces a competent stranger's email. The sharp twin produces something a colleague would recognise as yours: it opens with the point, it is short, it ends in an action with an owner, and it does not contain a single banned word. Nothing about the sharp version was thorough. It was specific, and that is what made it sound like you.

Practice check · not scored
Try the judgement call
Two sets of twin instructions. A: "You are a knowledgeable, professional CSM who writes clear and effective customer communications with a friendly and helpful tone." B: "I lead with the ask, never bury it. Max five sentences unless it is a QBR. I never say 'just checking in'. Every risk note ends with who does what by when. Write like the three samples below."

What specifically makes B behave like a real, identifiable person and A behave like a stranger?

The difference is specificity, not length or professionalism. B names concrete, personal choices (lead with the ask, five-sentence cap, banned phrases, a required ending, real writing samples) that distinguish this person from the average CSM, giving the model something real to imitate. A is made entirely of phrases true of almost everyone, so the model fills the gaps with the average and produces a stranger. Flexibility is not the goal, fidelity to you is.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. You want your twin to sound like you, but describing your own writing voice in words is hard. What is the most effective way to capture it?

Voice is far easier to show than to describe. Real examples of your writing carry your actual rhythm, vocabulary, and structure in a way that adjectives never capture, which is why pasting samples and saying 'write like this' works so much better than describing your tone. Generic instructions like 'write professionally' produce the average, and the model cannot infer your voice from your name.

2. A colleague's twin keeps producing bland, could-be-anyone output despite a long instruction set. Looking at their instructions, what is the most likely cause?

This is the generic-assistant trap. Long instructions can still be bland if every line is something true of most people ('professional, helpful, clear'). The model fills in the average because nothing distinguishes this person. The fix is specificity, the choices, phrases, and rules unique to them, not more length, a different tool, or a reset.

3. Why does this module insist on writing down 'what you would never do' as well as what you should do?

What you refuse to do is as much a part of your identity as what you do. Banned words and hard lines (never invent a stat, never say 'partnership', never send an unwritten apology) keep the twin from drifting into generic or off-brand output that you would never produce. It is not about looking thorough, the model handles positive instructions fine, and it is genuinely important, not optional.

Score: 0/3 ·

You will leave able to

  • Put your own way of working into clear words for the first time
  • Write instructions specific to you, using real examples of your voice
  • Avoid the generic-assistant trap that makes a twin sound like everyone else

Hands-on exercise

Write your twin brief: your role and standards, your voice (with three real anonymised writing samples), the checks you always run, and your hard lines and banned phrases. Keep it specific. Every generic line you leave in will come back as generic output.

The human element: The twin is a mirror of how well you understand your own craft. The act of writing it down often makes you a sharper CSM, because you finally name the standards you have been applying by instinct.

If you remember three things

  1. A digital twin is just an assistant set up to think and write like you, not like a generic stranger.
  2. Capture four things: your role and standards, your voice (shown with real examples), your standing checks, and your hard lines.
  3. Specific beats thorough. The parts of your description unique to you are the only parts that make the twin sound like you.
Part 1 · Build yourself · Module 04

Build a version of AI that works like you, part two

Giving your twin the knowledge it needs, in Claude, ChatGPT, Gemini, or Copilot, safely

12 min 85% to pass

The lesson

2halves of a twin

Last module you wrote down how you work. This module you build the actual assistant and give it what it needs to know: account context, your playbooks, product detail, past decisions. We show you how in each main tool, and where the hard line is on what should never be saved into an assistant.

1
Knowledge is what turns a voice into a twin

An assistant that sounds like you but knows nothing about your accounts is a good impressionist, not a twin. The knowledge layer is what lets it draft a renewal note that references the actual account, not a hypothetical one. The way every current tool does this is the same underneath: you give it documents, and when you ask a question, it retrieves the relevant bits and writes from them. That retrieval step is why how you organise the knowledge matters as much as what you put in.

The honest limit to know up front: the assistant does not read everything you give it on every question. It pulls the pieces that look relevant. So a tidy, well-labelled knowledge base gets good retrieval, and a giant unsorted dump gets hit-or-miss retrieval. Most twin failures are not the model being dim, they are the knowledge being badly organised.

The point

The model does not memorise your documents, it retrieves from them on each question. So the win is not stuffing in everything, it is organising what you put in so the right piece surfaces at the right moment. Structure beats volume.

2
The same build, in each tool

The principle is identical everywhere: a set of instructions (your twin brief from last module) plus a body of knowledge. Only the packaging differs. Pick whichever your organisation has approved and you already work in.

🟣
Claude Projects

A project holds your instructions plus uploaded knowledge, kept private to you or your team. Strong for analysis and writing. Nothing leaks between projects, so one project per account or workstream keeps context clean.

🟢
Custom GPT (ChatGPT)

Instructions plus up to around twenty knowledge files, plus optional actions that call external tools. Can be shared or published. Built-in web and code. Retrieval can be hit-or-miss, so label files clearly.

🔵
Gemini Gems

Instructions plus knowledge, with live Google Drive syncing and a very large context window. Strong if you live in Google Workspace, because it can read current Docs and Drive files rather than a frozen upload.

🟦
Copilot (M365)

Built into the apps you already use, and it reads the files and emails you point it at without the data leaving your organisation's tenant. The default home for anything touching real, un-anonymised account material.

The principle is portable

Instructions plus knowledge, every time. The tool is just packaging. Learn the principle and you can build your twin in whatever your company approves, and rebuild it in the next tool when things change.

3
The data line for a saved assistant

A one-off paste into a chat is a moment of exposure you control. A saved assistant is different: the knowledge sits there persistently, gets retrieved repeatedly, and may be shared. So the data rules from the foundational course apply harder here, not softer. What you load into a twin is a standing decision, not a one-time one.

The line: real customer personal data and commercially sensitive material only go into a tool your organisation has approved for that, inside its boundary, which for most people means Copilot in the M365 tenant or an enterprise agreement. For anything else, anonymise before it goes in, using the placeholder discipline. And never load security material, credentials, keys, or anything you could not defend being stored, because a saved assistant keeps it, retrieves it, and may surface it later in an answer you did not expect.

The standing-decision rule

Loading data into a saved assistant is a decision that keeps applying every time it answers, not a one-off paste. Approved tool and inside the boundary for real sensitive data, anonymise everything else first, and never load what you could not defend being stored and resurfaced.

The symptom

You load your twin with account notes, playbooks, and a year of QBR decks. You ask it for the renewal position on an account, and it confidently states a discount level that contradicts a note clearly sitting in its own knowledge base. The fact was right there. Why did it get it wrong?

Diagnosing it: structure or retrieval

Two possible causes. Either the knowledge is badly organised, so the contradicting note is buried in a 60-page dump where the wrong figure appears more prominently, or the retrieval simply did not surface the right document for that question. You test it by asking the assistant directly: "Which document are you drawing the discount figure from?" If it cites the wrong or a vague source, retrieval missed the right one. If it cites the right document but read it wrong, the document itself is ambiguous or buried.

The fix

Almost always the fix is structure, not a cleverer model. Split the giant dump into clearly named, single-topic documents ("Account X, current commercials", "Account X, renewal history"). Put the authoritative figure where it cannot be missed, and remove stale documents that contain outdated numbers. Well-organised knowledge gets reliable retrieval. A tidy library beats a smarter model almost every time.

Practice check · not scored
Try the judgement call
Your custom assistant, loaded with a single 80-page document containing everything about an account, gives a confident answer that contradicts a fact stated on page 60 of that same document.

What is the most likely cause and fix?

Assistants retrieve relevant pieces rather than reading everything every time, so a fact buried deep in an 80-page dump can simply fail to surface. The fix is structural: break the knowledge into clearly labelled, single-topic documents so retrieval can find the right piece. A more powerful model does not fix bad organisation, the page-60 fact is presumably correct, and uploading the same document repeatedly just clutters the knowledge base.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. You are about to build a twin that will need to reference real, un-anonymised customer account data, and you work in a Microsoft 365 organisation. Which is the most appropriate home for it?

When real un-anonymised customer data is involved, the data-residence question comes before the quality question, exactly as in the foundational course. In an M365 organisation, Copilot keeps the data inside the tenant boundary, making it the appropriate home. Choosing on writing quality alone ignores the data risk, a personal Custom GPT and a public Gem both put sensitive data outside the approved boundary.

2. Why does organising your twin's knowledge into clearly labelled, single-topic documents matter more than simply loading in everything you have?

The key mechanism is retrieval: the assistant pulls the pieces that look relevant to each question rather than reading the whole knowledge base every time. Well-structured, clearly named documents make the right piece surface reliably, a giant dump makes retrieval hit-or-miss. There is no one-document limit, the model does not read alphabetically, and it can open large files, it just may not retrieve the right part of them.

3. Why do the data rules apply more strictly to a saved assistant than to a one-off paste into a normal chat?

A one-off paste is a single, controlled moment. A saved assistant keeps the data, retrieves it repeatedly, and may surface it in future answers or to others it is shared with, so the decision to load something keeps applying. That standing nature is what raises the bar. It is not that saved assistants are inherently insecure or that they necessarily train on your data, it is the persistence and reuse that change the risk.

Score: 0/3 ·

You will leave able to

  • Build a real custom assistant in any of the four main tools
  • Organise its knowledge so the right piece actually surfaces when asked
  • Apply the data line to a saved assistant, treating loading as a standing decision

Hands-on exercise

Build your twin in your approved tool. Load your instructions from last module, then add three or four clearly named, single-topic knowledge documents (anonymised unless the tool is approved for real data). Test it with five real questions and check it retrieves the right facts. Where it misses, fix the structure, not the model.

The human element: A twin is a powerful convenience, and convenience is exactly where data discipline slips. The moment something becomes effortless to reach is the moment to be most careful about what you put in it.

If you remember three things

  1. Knowledge is what turns a voice into a twin, and every tool does it the same way underneath: instructions plus knowledge, retrieved per question.
  2. Structure beats volume. Clearly named single-topic documents get reliable retrieval, big dumps do not.
  3. Loading data into a saved assistant is a standing decision: approved tool for real sensitive data, anonymise the rest, never load what you could not defend being stored.
Part 1 · Build yourself · Module 05

Write instructions that actually hold up

Why assistants drift, and how to write rules that stay steady under pressure

12 min 85% to pass

The lesson

9out of 10 is not enough

Most custom assistants drift. They ignore their own rules, behave differently each time, or fall apart on an odd request. This module is about writing instructions that hold steady even under pressure, the difference between a fun toy and something you would trust in front of a customer.

1
Why assistants drift

You build a twin, test it on a few normal requests, it behaves, and you trust it. Then three weeks later it does something off: uses a banned phrase, invents a figure, answers a question it should have refused. It did not break. It drifted, because your instructions held for the easy cases you tested and not for the awkward one you did not.

The reason is simple. A vague instruction is a suggestion, and the model follows suggestions most of the time, which is exactly the problem. Most of the time is fine for a toy and not fine for customer-facing work. The job of this module is turning suggestions into rules that hold on the tenth try, not just the first nine.

The point

An instruction that works nine times out of ten is not 90 per cent good, it is a problem waiting for the tenth customer. Reliability under the awkward case, not the easy case, is what separates a real tool from a toy.

2
The three things that make instructions hold

Holding instructions are built, not wished for. Three techniques do almost all the work, and they are the same techniques whether you are in Claude, ChatGPT, Gemini, or Copilot.

🔒
Hard constraints, not soft hopes

"Try to be concise" is a hope. "Never exceed five sentences. If you cannot fit it, say so and stop" is a rule. State limits as absolutes with a defined behaviour when they are hit, not as preferences the model can weigh away.

📐
Worked examples

Show, do not just tell. One example of a good output and one of a bad output, with a note on why, teaches the model more than a paragraph of description. Examples pin behaviour the way nothing else does, especially for voice and format.

🚫
Explicit refusals

Name what it must not do and what to do instead. "If asked to state a figure not in your knowledge, do not estimate. Reply: I do not have that figure, here is what I would need." An assistant that knows how to refuse is far safer than one that always tries to help.

The craft

Absolutes over preferences, examples over descriptions, and a defined refusal for the cases that matter. These three turn an assistant that mostly behaves into one that behaves even when the input is awkward.

3
Test it like software, not like a hope

The mistake is testing your assistant only on the requests you expect. Of course it handles those, you built it for them. The behaviour that bites you lives in the awkward inputs: the question that is slightly outside its job, the request that tempts it to invent, the message phrased to push past a rule. You have to go looking for those on purpose.

So test adversarially. Try to make it break. Ask for the figure it should not know. Phrase a request to tempt a banned word. Give it a messy, half-relevant input and see if it stays in its lane. Each time it slips, you have found a missing constraint, and you fix the instruction, not just that one answer. A few rounds of trying to break it is worth more than fifty normal uses, because it finds the tenth-try failure before a customer does.

What you walk away with

A hardened, tested instruction set: absolutes, worked examples, and refusals, proven against inputs designed to break it. The goal is not an assistant that works when you are gentle with it. It is one that holds when reality is not.

The slip

Your twin is told "be accurate and do not make things up". In testing it behaves. Then someone asks it "what was their adoption rate last quarter?" and the figure is not in its knowledge. Instead of saying so, it produces a confident, plausible number. The instruction "do not make things up" was a hope, and the model's stronger instinct to be helpful won.

Finding the missing constraint

The failure is not random, it is the predictable result of a soft instruction meeting a tempting gap. "Do not make things up" tells the model what not to do but not what to do instead, so when it lacks the figure it falls back on its default: produce a helpful-sounding answer. The gap in the instruction is the absence of a defined behaviour for missing information.

The fix that holds

Replace the hope with a hard refusal and a defined fallback: "If asked for any specific figure, date, or name not present in your knowledge, do not estimate or infer it. Reply exactly: I do not have that in my knowledge. Here is what I would need to answer." Then test it again on the same question, and on three more phrased to tempt invention. Now the awkward case has a rule, and the assistant holds where it used to slip.

Practice check · not scored
Try the judgement call
An assistant is instructed: "Always write in a warm, concise tone and avoid jargon." In testing it is great. In real use it occasionally produces long, jargon-filled replies when the input is complex.

What is the underlying problem?

"Always" and "avoid" sound firm but are still preferences the model balances against other goals, and complex input tips the balance. The fix is a hard constraint with a defined behaviour: "Never exceed [N] sentences. Never use these banned terms. If the topic is complex, say it is complex and ask what to prioritise." The model is not faulty, the two qualities are compatible, and repetition of a soft instruction does not make it hold.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Your assistant follows its rules on nine normal requests and breaks on the tenth, an unusual one. What is the right way to think about this?

For customer-facing work, the awkward tenth case is exactly the one that bites, so 'good enough' is the wrong frame. The slip points to a missing or soft constraint, and the durable fix is to harden the instruction so the rule holds under that kind of input, then re-test. Fixing only the single answer leaves the gap open for next time, and a full rebuild is overkill when one constraint is missing.

2. What is the single most reliable way to pin down an assistant's voice and output format so they stay consistent?

Examples pin behaviour far more reliably than description, especially for voice and format, because they show the target rather than approximate it in words. A good example plus a bad example with the reason teaches the model the boundary. Longer descriptions still leave room for interpretation, 'be consistent' is a vague instruction, and turning down creativity changes randomness, not whether the model knows what good looks like.

3. Why does the module insist on testing your assistant with inputs deliberately designed to break it, rather than just normal requests?

Testing only the requests you expect confirms what you already built for and misses the tenth-try failures. Adversarial testing, trying to tempt invention, push past a rule, or knock it out of its lane, surfaces the missing constraints before a customer-facing situation does. It does not make the model smarter, and it is meaningfully better than normal testing precisely because it targets the cases that actually break things.

Score: 0/3 ·

You will leave able to

  • Write instructions that survive messy and difficult requests, not just easy ones
  • Use absolutes, worked examples, and explicit refusals to keep behaviour steady
  • Test an assistant adversarially instead of hoping it behaves

Hands-on exercise

Take the twin you built and try to break it. Ask for a figure it should not know, phrase a request to tempt a banned word, feed it a messy off-topic input. Each time it slips, add a hard constraint, an example, or a refusal to fix it. Stop when a few rounds of trying to break it stop working.

The human element: A reliable assistant is built by someone willing to attack their own work. The instinct to test gently, to confirm it works, is the instinct that ships the tenth-try failure to a customer. Be your own adversary first.

If you remember three things

  1. Assistants drift because soft instructions hold for easy cases and fail on awkward ones.
  2. Three things make instructions hold: hard constraints not soft hopes, worked examples not descriptions, and explicit refusals with a defined fallback.
  3. Test adversarially. Go looking for the tenth-try failure on purpose, and fix the instruction, not just the answer.
Part 2 · Build systems · Module 06

From an assistant that answers to an agent that acts

What an agent really is, where the builders are today, and guardrails before power

12 min 85% to pass

The lesson

answer→ act

An assistant answers questions. An agent goes and does things: it works towards a goal, takes several steps, and takes real actions. That is a big jump, and it comes with real risk. This module is about what an agent actually is, and putting the guardrails in before the power.

1
What an agent actually is

The word "agent" is overused, so here is the honest definition. An assistant responds: you ask, it answers, you decide what to do. An agent is given a goal and then takes steps towards it on its own, including actions in the world: reading data, calling tools, sending things, updating records. The shift is from a thing that produces words to a thing that produces actions.

That shift is the whole reason agents are powerful and the whole reason they are dangerous. A bad answer from an assistant is a paragraph you can ignore. A bad action from an agent is something that actually happened: an email that went out, a record that changed, a customer who got contacted. So the entire discipline of building agents is defining the goal tightly and deciding, in advance, where the agent must stop and ask a human.

The point

An agent does not just answer, it acts. That is its power and its danger in one sentence. Everything about building one safely follows from taking the 'it acts' part seriously.

2
Define an agent by its goal and its limits, not its steps

Beginners describe an agent as a list of steps: do this, then this, then this. That is brittle, because reality rarely matches the script, and a step-following agent breaks or improvises badly the moment the input is unexpected. The better way is to define two things: the goal (what done looks like) and the limits (what it must never do, and where it must stop and hand to a human).

Think of it like briefing a capable but new team member on a task that touches customers. You would not just list steps, you would say "here is what success looks like, here are the lines you do not cross, and here are the points where you check with me before acting". An agent needs exactly that: a clear goal, hard limits, and defined human checkpoints. The steps it can largely work out. The limits and checkpoints are what keep it safe.

Goal and guardrails

Define the agent by what done looks like and what it must never do, plus where it stops for a human. Defining it by a rigid list of steps is brittle and unsafe, because the danger is in the unexpected input the script did not cover.

3
Where the human checkpoint goes

The most important design decision in any agent is where it stops and waits for you. The rule is the one from the foundational course, applied to action: automate assembly, never judgement. The agent can freely do the gathering, sorting, drafting, and formatting. It must stop before anything that is a judgement call or an irreversible action: sending a customer message, changing a record that matters, committing to anything, or deciding which risk is real.

The checkpoint sits between judgement and action
Safe to automate
Assembly
  • Gather data
  • Sort and analyse
  • Draft messages
  • Surface patterns
  • Propose options
Stophuman review
Never autonomous
Action
  • Send customer messages
  • Change records that matter
  • Commit on your behalf
  • Decide which risk is real
  • Anything irreversible

Build agents that propose, not execute. Most of the value, little of the risk.

Today's no-code and low-code builders make this easier than it used to be: you can build agents inside the major assistant platforms, connect them to tools and data, and crucially set them to propose rather than execute. The honest state of play is that a CS agent that drafts and proposes is safe and genuinely useful right now, while one that acts on customers entirely unsupervised is not something to trust yet. Build the proposing kind, put the checkpoint before every real action, and you get most of the value with little of the risk.

What you walk away with

A clear plan for a small, safe agent: a defined goal, hard limits, and a human checkpoint before every judgement call or irreversible action. Plus a first build that proposes rather than executes, which is where the safe value is today.

What was built

A CSM builds an agent with the goal "keep my at-risk accounts engaged" and lets it both identify at-risk accounts and send a check-in email. It runs overnight. In the morning, a healthy account that happened to have one quiet week has received an oddly worried check-in email, and the customer replies asking if something is wrong.

Was the goal wrong, or the guardrail missing?

The goal was reasonable, keeping at-risk accounts engaged is a fine aim. The failure was a missing guardrail: the agent was allowed to take an irreversible, customer-facing action (sending an email) based on its own judgement (deciding the account was at risk) with no human checkpoint between the two. The single quiet week was a weak signal the agent over-read, and because it could act, the misread became an email a customer actually received.

Where it should have stopped

The checkpoint belongs exactly between judgement and action. The agent should have done all the assembly, identified the possibly-at-risk accounts, and drafted the check-in emails, then stopped and presented them to the CSM: "these five accounts look quiet, here are draft check-ins, send which ones?" The CSM would have instantly seen the healthy account, removed it, and sent the rest. Same time saved, none of the risk. The lesson: the agent could assemble and propose freely, it just should never have crossed into sending on its own.

Practice check · not scored
Try the judgement call
An agent is given the goal "reduce my overdue admin" and the ability to update CRM records and close out tasks. It runs and marks a batch of genuinely open tasks as complete because they were old, distorting the team's reporting.

What was the core design failure?

The goal was fine; the failure was letting the agent act on a judgement ('old equals done') by changing records, with nothing stopping it for human review. Closing tasks is consequential and distorts reporting, exactly the kind of action that needs a checkpoint. It is not that agents can never touch a CRM, assembly and drafting are fine, it is that the judgement-to-action step needed a human. A faster model would just have made the wrong changes sooner.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. What is the single most important thing that makes building an agent different from building an assistant?

The defining difference is action. An assistant produces words you can choose to use or ignore, an agent takes steps in the world, so a mistake becomes an email sent or a record changed. That is why agent design centres on goals, limits, and human checkpoints. It is not about model choice, instruction length, or cost, those are incidental next to the fact that an agent acts.

2. Why is defining an agent by a rigid list of steps a worse approach than defining it by a goal plus limits?

Reality rarely matches a script, so a step-following agent fails or improvises poorly the moment something unexpected appears, and unexpected inputs are exactly where the risk lives. Defining the goal (what done looks like) and the limits (what it must never do, where to stop) lets the agent handle variation while staying safe. The difference is real and is about robustness, not effort or how easily instructions are followed.

3. Given the honest state of CS agents today, what is the safe and useful way to build one?

The realistic sweet spot today is an agent that does all the assembly and drafting then stops and proposes, with a human approving any real action. That captures most of the value with little of the risk. Full autonomy on customers is not trustworthy yet, avoiding agents entirely throws away genuine value, and reviewing only after the fact means the irreversible actions have already happened, which is too late for the cases that matter.

Score: 0/3 ·

You will leave able to

  • Describe an agent by its goal and limits, not a brittle list of steps
  • Pick the right builder and set it to propose rather than execute where it matters
  • Put the human checkpoint exactly where it protects you and the customer

Hands-on exercise

Pick one repeatable process you do. Write its agent design on one page: the goal (what done looks like), the hard limits (what it must never do), and the exact points where it must stop and hand to you. Mark which steps are assembly (safe to automate) and which are judgement or action (needs a checkpoint). Do not build it yet, just design it safely.

The human element: An agent's guardrails are a statement of your judgement, not a limitation on the tool. Deciding where it must stop is the most senior thing you do in the whole build, because it is where you take responsibility for what it does in your name.

If you remember three things

  1. An agent acts, an assistant answers. That difference is the source of both its power and its danger.
  2. Define an agent by its goal and limits, not a rigid list of steps, because the danger is in the unexpected input.
  3. Automate assembly, never judgement. Put a human checkpoint before every customer-facing or irreversible action, and build agents that propose rather than execute.
Part 2 · Build systems · Module 07

Build a real CS agent from start to finish

One genuinely useful agent, from idea to tested tool you could defend to your VP

12 min 85% to pass

The lesson

Honest expectation

Most CSMs who take this module will not build a fully autonomous agent, and they shouldn't. The point is to know exactly how far you could go, then make a deliberate choice about how far to go. A custom GPT plus your prompt library is a complete answer for most working CSMs. What you build here is a narrow, useful, scoped agent (one or two of them, not a multi-step autonomous system). The skill is in the design, not the ambition.

1narrow agent, built well

This is the big one. You take one genuinely useful, narrow agent all the way from idea to working tool: maybe one that watches your portfolio for warning signs and drafts an alert, or one that pulls together a call brief from several sources. With testing, failure handling, and the standard where you could defend it to your VP.

1
Pick one agent worth building

The temptation is to build something ambitious. Resist it. The right first agent is narrow, useful, and run often: a portfolio-watch agent that scans your accounts for risk and expansion signals and drafts you a ranked Monday brief, or a call-prep agent that assembles a full brief from your notes, the CRM, and public information. Both are mostly assembly, both run every week, and both stop at a clear point for your judgement. That makes them safe to build and genuinely worth the effort.

Before any building, write the design from the last module: the goal, the limits, and the checkpoints. The portfolio-watch agent's goal is "a ranked weekly watchlist with draft alerts". Its hard limit is "never contact a customer, never change a record". Its checkpoint is "present the watchlist and drafts to me, I decide what happens". That one page is the spec you build against, and it is what keeps the build honest.

The point

Your first real agent should be narrow, frequent, and mostly assembly, with a clean stopping point for your judgement. Ambition is the enemy of a working first build. Pick the boring, weekly, useful one.

2
Build it, then break it

Building the happy path is the easy 20 per cent: connect the data, run the analysis, format the output, stop at the checkpoint. The real work is the other 80 per cent, handling the ways it goes wrong, because an agent that only works on clean inputs is a demo, not a tool.

So once it runs on good data, deliberately feed it bad data, the same adversarial habit from module 5, now applied to a thing that acts. Give it an account with missing fields. Give it a messy export with a column out of place. Give it an account that is genuinely ambiguous. Watch what it does. Does it flag the gap, or invent around it? Does it stop at the checkpoint, or quietly do something? Each failure you find is a limit you add or a checkpoint you reinforce before the agent ever runs on your real book.

The 80 per cent

Building the happy path is the easy part. The work is failure handling: feed it messy, missing, and ambiguous data on purpose and fix what breaks. An agent that only works on clean inputs is a demo. One that fails safely on bad inputs is a tool.

3
Write it up so it survives you

An agent that lives only in your head is a personal trick, and personal tricks do not become authority or career capital. The final step is documenting it well enough that a colleague could pick it up: what it does, what it must never do, where it stops, how to run it, and what its known weaknesses are. This is the same discipline as a good library entry, scaled up to a working system.

The documentation does double duty. It is what lets you hand the agent to a teammate, which is how one good build becomes a team asset (and how you become the person who built it). And it is what lets you defend the agent when your VP, or security, asks "what is this thing, what can it touch, and how do you know it is safe?" An agent you can explain on one page is one you can defend. An agent you cannot is a risk wearing a productivity costume.

What you walk away with

One complete, working CS agent you will actually use, tested against bad inputs, and written up clearly enough to hand to a colleague or defend to your VP. That write-up is what turns a clever build into a real asset.

The failure

Your portfolio-watch agent works perfectly in testing. You point it at your real book and one account comes back marked low risk and quietly dropped to the bottom of the list, when you happen to know it is in trouble. The agent did not flag it. On a real account, with real consequences, it missed.

Tracing it to the real cause

You do not guess, you trace. You look at what data the agent actually received for that account and find the cause: the account's recent activity was in a field the agent was not reading, or the export had that account's key column blank, so the agent saw no signals and concluded no risk. The model did exactly what it was told with the data it got. The failure was upstream, in the data the agent could see, not in its reasoning.

The fix and the lesson

Two fixes. First, the immediate one: make the agent read the field it was missing, or handle the blank column by flagging "insufficient data on this account" rather than silently scoring it low. That last point is the real lesson: an agent should never treat missing data as good news. Absence of signal is not absence of risk. You add a hard rule, "if key data is missing for an account, flag it for manual review rather than scoring it", and you have turned a silent failure into a safe one. This is exactly the kind of fix you only find by testing on messy real data, not clean test data.

Practice check · not scored
Try the judgement call
Your call-prep agent assembles great briefs in testing. On a real account it produces a brief that confidently states the renewal date as a date you know is wrong. You trace it and find the CRM export it read had that account's renewal date field empty.

What is the right fix?

The agent invented a date because nothing told it what to do with a missing field, so it filled the gap, the confident-invention failure applied to an agent. The fix is a hard rule for missing data: flag it as missing and surface the gap, never fill or guess. Accepting it ships wrong dates to calls, abandoning the agent throws away a working tool over a fixable gap, and a 'better' model guessing more plausibly is worse, not better, because a convincing wrong date is more dangerous than an obvious one.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. What makes a good candidate for your first real CS agent?

The right first agent is narrow, frequent, mostly assembly, and stops cleanly for your judgement, which is why portfolio-watch or call-prep work so well. Ambition makes a first build fail. A rare task fails the frequency test from module 1. And letting the agent contact customers directly is exactly the unsupervised-action risk the previous module warned against, not a feature to chase.

2. You have built an agent that works on clean test data. Why is that only the easy 20 per cent of the job?

An agent that only handles clean inputs is a demo, because real account data is messy, incomplete, and sometimes ambiguous. The bulk of the work is deliberately feeding it bad data and fixing how it fails, so it flags gaps rather than inventing around them. Documentation matters but is a separate step, clean-data testing does prove the happy path works, and visual polish is not the point, safe failure is.

3. Your agent silently scored an account as low risk because the data it needed was missing. What is the right principle to build in?

Absence of signal is not absence of risk, so an agent must never read missing data as good news. The safe principle is to flag incomplete accounts for manual review rather than scoring them. Filling gaps with estimates reintroduces confident invention, silently excluding them hides accounts that may be the most at risk, and assuming missing means fine is the exact silent failure the example warned about.

Score: 0/3 ·

You will leave able to

  • Take an agent from idea to a tested, working tool
  • Handle the ways it fails on messy real data, not just the happy path
  • Write it up so it becomes a team asset you can hand over or defend

Hands-on exercise

Build the agent you designed last module, narrow and useful. Get the happy path working, then spend most of your time breaking it: feed it accounts with missing fields, messy exports, and ambiguous cases. Add a rule for every failure, especially 'flag missing data, never score around it'. Then write the one-page doc: what it does, what it must never do, where it stops, and its known weaknesses.

The human element: The agent you can defend on one page is the one worth having. If you cannot explain what it touches and how you know it is safe, you do not yet own it, it owns a piece of your risk.

If you remember three things

  1. Pick a narrow, frequent, assembly-heavy agent with a clean stopping point. Ambition kills first builds.
  2. The happy path is the easy 20 per cent. The work is failure handling: break it with bad data on purpose and fix what slips.
  3. Never let an agent treat missing data as good news, and write it up so you can hand it over and defend it.
Part 2 · Build systems · Module 08

Judge AI output like an expert

Catching confident-but-wrong output, and the honest truth about when to trust the tool over your gut

14 min 85% to pass

The lesson

looks right≠ is right

The thinking heart of the course, and where we finally go into the nuance the foundational course held back. How do you tell genuinely good output from confidently wrong output? And the honest question: good tools really do weigh several signals together, so when is the tool's read better than your gut, and when is overruling it just ego?

1
The skill that matters most as you build more

The more you build, the more output flows past you, and the more dangerous it becomes to wave it through because it reads well. Fluent, confident, well-structured output is exactly what the model is best at producing, whether or not it is correct. So the senior skill, the one this whole course builds towards, is judging output on its substance, not its polish.

The foundational course taught a simplified rule: when signals stack up, trust your judgement and override the tool. That was the right rule for a beginner, because the common failure was deferring to a comfortable score and getting blindsided. But it was a simplification, and now you are ready for the honest, harder version: sometimes the tool's read genuinely is better than yours, and knowing which case you are in is the actual expert skill.

The point

As you build more, your job shifts from producing output to judging it. And the model is best at producing output that looks right, regardless of whether it is. Judging substance over polish is the skill that protects everything you have built.

2
Catching confident-but-wrong output

Three failure types hide inside fluent output, and a quick structured check catches most of them. This is the checklist you run before acting on anything that matters.

👻
Quiet invention

A fact, figure, or name that was never in your input, sitting in a fluent sentence next to true ones. Check: can I trace every specific claim back to something I actually provided? Anything I cannot trace is suspect until verified.

🔗
Broken reasoning

A conclusion that does not actually follow from the steps, or a step quietly skipped. Check: does each step lead to the next, and does the conclusion really follow, or does it just sound conclusive? Read the logic, not the confidence.

🎯
Plausible-but-wrong framing

The answer is internally consistent but built on a wrong assumption it never states. Check: what is this assuming that I did not tell it, and is that assumption actually true? The most dangerous errors are the unstated ones.

The check

Trace every specific claim to a source, read the reasoning rather than the confidence, and surface the unstated assumptions. A few minutes of this on anything that matters catches the errors that fluent prose hides. Run it as a habit, not a one-off.

3
When to trust the tool, and when to overrule it

Here is the honest nuance. Modern tools genuinely do weigh many signals together, and on a question that is mostly about combining lots of data points consistently, the tool's read can be better than your gut, your gut is moody, tired, and biased by the last account that burned you. Overruling it there is ego, not judgement. But on a question that turns on context the tool does not have, what your champion's tone really meant, what is happening politically that is in no dataset, your read wins, and deferring to the tool is the lazy move.

The test question that tells them apart
What does the right answer here depend on most?
If: combining data

Lean tool, check your ego

Many signals weighed consistently is the tool's strength. Your gut is moody and biased by the last burn. Overruling here is ego.

If: context only you have

Lean judgement, hold ground

A tone, a private exchange, an unstated political shift, is in no dataset. Deferring to a confident score is the lazy move.

Same shape, opposite answers. The test question is the whole skill.

So the expert test is not "trust" or "overrule" as a fixed stance. It is one question: does the right answer here depend mainly on combining available data, or mainly on context only I have? If it is the data, lean towards the tool and check your urge to override. If it is the context, lean towards your judgement and do not let a confident score talk you out of what you know. The skill is telling the two situations apart, because they can look almost identical on the surface.

What you walk away with

A short checklist for sense-checking any output, plus the real test for trust: does this turn on combining data the tool has, or on context only you have? Data, lean tool and check your ego. Context, lean judgement and hold your ground.

Case one: trust the tool

Your portfolio tool ranks an account as higher risk than you would have. Your gut says it is fine, you had a good call last month. But when you look, the tool is weighing twelve signals consistently: usage trend, support volume, login breadth, contract timing, and more, across the whole quarter. Your 'fine' is based on one good call and a general feeling. This question is mostly about combining many data points, which is exactly what the tool does better than a busy human. Overruling it here would be ego. The right move is to trust the read and investigate, your gut was the weaker instrument.

Case two: overrule the tool

A near-identical situation: the tool ranks an account low risk, the data looks healthy. But you know something the data does not, on last week's call the champion mentioned, almost in passing, that they are moving to a competitor's parent company in a reorg. That single piece of context outweighs all twelve healthy signals, and it is in no dataset. Here the right answer turns on context only you have, so your judgement wins decisively. Deferring to the healthy score because it is confident would be the lazy failure.

Why they look identical and are not

On the surface both are 'tool says one thing, I think another'. The difference is invisible until you ask the test question. Case one turns on combining lots of data, the tool's strength, so you check your ego and trust it. Case two turns on a piece of context the tool cannot have, your strength, so you hold your ground. Same shape, opposite answers, and the only thing that tells them apart is asking what the right answer actually depends on. That question is the whole skill.

Practice check · not scored
Try the judgement call
Your tool flags an account as high risk based on a broad, consistent weighing of usage, support, and engagement data across the quarter. Your instinct says it is fine, based mainly on a good vibe from one recent call. You are tempted to override the tool.

What is the right move and why?

This is the 'trust the tool' case. The decision rests mainly on combining many signals consistently across a quarter, exactly the tool's strength, while your instinct rests on a single call and a feeling, which is the weaker instrument here. Overriding would be ego, not judgement. 'My judgement always wins' is the beginner oversimplification this module corrects, one good call is not the strongest signal against twelve weighed ones, and waiting dodges the decision the test is asking you to make.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. The foundational course said 'when signals stack up, trust your judgement and override the tool'. Why does this advanced module call that a simplification?

The foundational rule was right for beginners, whose common failure was deferring to a comfortable score, but it was a simplification. The honest version is that the tool can be the better instrument when a decision turns on combining many data points consistently, so overruling it there is ego. That does not mean always trust the tool, that the old advice was wrong, or that human judgement is obsolete, it means knowing which kind of question you are facing.

2. What is the actual test for whether to trust the tool's read or your own judgement on a given account?

The expert test is to ask what the right answer actually depends on. If it hinges on combining lots of available data consistently, that is the tool's strength, so lean towards it and check your urge to override. If it hinges on context the tool cannot have, like a champion's offhand remark, that is your strength, so hold your ground. The tool's confidence is not evidence, personal liking is bias, and seniority is not the test.

3. When sense-checking output that matters, which is the most dangerous kind of error to miss?

Plausible-but-wrong framing, an answer built on an unstated assumption that happens to be false, is the most dangerous because everything downstream looks coherent and confident while resting on a broken foundation. That is why the check 'what is this assuming that I did not tell it, and is it true?' matters so much. Spelling, length, and formatting are cosmetic by comparison and do not make a correct-looking answer actually wrong.

Score: 0/3 ·

You will leave able to

  • Run a quick check that catches invention, broken reasoning, and false assumptions
  • Say clearly when the tool's read beats your gut, and when it does not
  • Avoid both blindly trusting output and overruling it out of habit

Hands-on exercise

Take a real AI output you acted on recently. Run the three checks: trace every specific claim to a source, read whether the reasoning actually holds, and name the unstated assumptions. Then take a recent account call where you and a tool disagreed, and ask the test question: did it turn on combining data, or on context only you had? Notice which way you should have leaned.

The human element: The honest expert is the one who can say 'the tool was right and I was about to let my ego override it' as easily as 'I knew something the data could not'. Both humility and confidence, applied to the right cases, are the skill. Neither alone is.

If you remember three things

  1. As you build more, judging output on substance over polish becomes your most important skill.
  2. Catch the three hidden failures: quiet invention, broken reasoning, and false unstated assumptions.
  3. The trust test: does the answer turn on combining data (lean tool, check your ego) or on context only you have (lean judgement, hold firm)?
Part 2 · Build systems · Module 09

Join the pieces into something that runs itself

Chaining your builds into a reliable routine, with checks so one failure does not cascade

14 min 85% to pass

The lesson

8good, 1 damaging

Now you connect what you have built. One step hands off to the next, an assistant feeds an agent, and a whole routine runs with little input from you. We are honest about where chains break and how to keep them steady, with checks and backups so one failure does not knock everything over.

1
What orchestration actually buys you

So far you have built separate pieces: prompts, a twin, an agent. Orchestration is connecting them so the output of one becomes the input of the next, and a whole routine runs as one flow. The Monday morning system that pulls your exports, runs the portfolio scan, drafts the watchlist, and queues your call preps, before you have finished your coffee, is orchestration. Done well, it compounds: the time saving is not one task but a whole chain of them.

But chaining adds a new kind of risk that single tools do not have. When one tool gives a slightly wrong output, you usually catch it. When tool A's slightly wrong output silently becomes tool B's input, the error travels and grows, and by the end of the chain you have a confident, wrong result with no obvious sign of where it went wrong. The power of orchestration and its danger are the same thing: the steps feed each other.

The point

Orchestration connects your builds so a whole routine runs as one flow, which compounds the time saving. The same connection means an error in one step silently feeds the next, so the danger compounds too. Reliability, not cleverness, is the goal.

2
Where chains break, and how to hold them

Chains fail at the joins, where one step's output becomes the next step's input. Three disciplines keep them steady, and they are worth more than any amount of clever connecting.

Errors enter once, then compound
1
Gather and clean

Pull data, normalise it

checkpoint here
2
Analyse

Score, classify, surface patterns

3
Draft

Turn analysis into a polished output

The most valuable checkpoint sits at the earliest join where bad input can enter. Catch it there before it polishes itself into a confident wrong answer.

🚦
Checkpoints at the joins

Do not let a step feed the next one blindly on anything that matters. Put a check at the join: a quick validation, or a human glance, before the output flows on. The cost of a checkpoint is seconds. The cost of a silent error travelling the whole chain is a wrong result you trust.

🪂
Fallbacks for missing inputs

Decide in advance what each step does when its input is missing or malformed. Stop and flag, rather than carry on with a guess. A chain that pauses safely beats one that completes confidently on bad data.

🔍
A way to see where it failed

When the end result is wrong, you need to find which step caused it. Keep each step's output visible, not hidden inside the chain, so you can trace the failure to its source instead of guessing. A chain you cannot inspect is a chain you cannot trust.

The craft

Checkpoints at the joins, defined fallbacks for missing inputs, and visibility into each step. These turn a clever chain that works most of the time into a reliable one that fails safely and tells you where.

3
When to chain, and when not to

Not everything should be chained. Every join you add is a place that can break and a place an error can hide, so chaining has a real cost in fragility, not just setup. The honest rule: chain steps that are stable, mechanical, and run often, and keep a human in the loop at any point where judgement matters or an action touches a customer.

The best orchestrated systems are mostly assembly with judgement checkpoints, not fully automated end to end. The Monday system that assembles everything and then stops for you to make the calls is better than one that tries to decide and act for you, because the assembly is where the reliable time saving is, and the judgement is where the risk is. Chain the boring parts. Hold the deciding parts. A routine that does 90 per cent of the gathering and hands you a clean decision point is the sweet spot, more automation than that usually buys fragility, not value.

What you walk away with

A connected routine linking at least two of your builds, with checkpoints at the joins, fallbacks for bad inputs, and a human holding the judgement calls. Chain the boring, stable, frequent parts. Never chain past a point where judgement or a customer action lives.

The setup

You chain three things: an agent pulls and cleans your weekly account exports, your twin analyses them for risk, and a drafting step produces ranked alert emails. For eight weeks it is brilliant, a full Monday brief assembled before you sit down. On the ninth week, it produces a confident brief that badly misranks your accounts, and you nearly act on it.

Tracing where it broke

Because you kept each step's output visible, you can trace it. The export step ran on a week where the source system had changed a column, so the cleaned data was subtly wrong, a date column had shifted. The analysis step received this malformed input and, having no fallback for it, analysed it anyway and produced a confident but wrong ranking. The drafting step faithfully turned the wrong ranking into a polished brief. The error entered at step one and grew, invisible, through steps two and three.

Where the check goes, and why there

The checkpoint belongs at the first join, right after the export-and-clean step, because that is where bad data enters and where catching it stops the error before it travels. A simple validation there, does this data look like last week's shape, are the key columns present and sane, would have caught the shifted column and paused the chain with a flag, rather than letting the malformed data flow into analysis. Putting the check at the end (reading the final brief critically) helps, but it is later and relies on you spotting a polished wrong answer. The earliest join where bad input can enter is almost always where the most valuable checkpoint goes: catch the error at its source, before it compounds.

Practice check · not scored
Try the judgement call
A three-step chain (gather data, analyse it, draft the output) gives excellent results most weeks, then occasionally produces a confidently wrong final output. You can only add one checkpoint.

Where should it go and why?

Errors in a chain compound from where they enter, so the most valuable single checkpoint is at the earliest join where bad input can appear, here right after data gathering. Catching malformed data there stops it before analysis and drafting build confident wrong output on top of it. An end checkpoint catches it only after it has travelled and been polished, position genuinely matters so 'anywhere is equal' is wrong, and the middle step is not special just because it is the cleverest.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. What new kind of risk does chaining tools together introduce that single tools do not have?

The defining risk of orchestration is error propagation: because each step feeds the next, a small mistake in an early step becomes the input to later steps and compounds into a confident, hard-to-spot wrong result. That is precisely why checkpoints at the joins matter. Speed is not the issue, chaining does not change the underlying model, and 'several safe tools in a row' misses that the joins themselves are where the new risk lives.

2. Why is it important to keep each step's output visible rather than hidden inside the chain?

A chain you cannot inspect is a chain you cannot trust, because when something goes wrong you have no way to find the source. Keeping each step's output visible lets you trace a bad final result back to the step that caused it, as in the shifted-column example. It is not about speed or the model reading itself, and visibility stays important precisely because chains that work today can break later in ways you will need to diagnose.

3. What is the honest sweet spot for an orchestrated CS routine?

The reliable value is in chaining the boring, stable, frequent assembly work, while the risk lives in judgement and customer-facing actions, which is why those keep a human. A routine that assembles 90 per cent and hands you a clean decision point is the sweet spot. Full end-to-end automation buys fragility past that point, avoiding chaining entirely throws away real value, and there is no reason data-gathering specifically must be manual, it is often the safest part to automate.

Score: 0/3 ·

You will leave able to

  • Join steps into a routine that is reliable, not just clever
  • Add checkpoints and fallbacks so one failure does not cascade
  • Judge when chaining is worth the fragility it adds, and when it is not

Hands-on exercise

Connect two things you built into a small chain (for example, your portfolio agent feeding your twin's analysis). Then deliberately break it: feed the first step malformed data and watch what reaches the end. Add a checkpoint at the first join that validates the input before it flows on, and confirm the chain now pauses safely instead of producing a confident wrong result.

The human element: A system that runs itself is seductive, and the seduction is exactly where judgement quietly leaks out. The best orchestrators automate the assembly ruthlessly and guard the decision points jealously. The machine sets the table. You still decide what to do at it.

If you remember three things

  1. Orchestration compounds the time saving by connecting your builds, and compounds the risk because errors feed forward.
  2. Hold chains together with checkpoints at the joins, fallbacks for bad inputs, and visibility into each step.
  3. Chain the stable, mechanical, frequent parts. Keep a human at every point of judgement or customer action.
Part 3 · Lead · Module 10

Become the AI person for your whole company

Building a shared CS knowledge base, and owning the accuracy that makes you the authority

12 min 85% to pass

The lesson

1wrong answer to everyone

The step from helping yourself to helping everyone. You build a shared knowledge base that an assistant can use for the whole team, and you take on the job of keeping it accurate, which is what makes you the go-to person. The trap is honest: one wrong answer given to the whole team is worse than no answer at all.

1
From personal tool to company capability

Everything so far has made you faster. This module is where you become valuable to the whole organisation, by building something the team uses, not just you. A shared CS knowledge base, an assistant the whole team can ask, that knows your products, your playbooks, your standard answers, turns your personal fluency into a company capability. And the person who builds and owns it becomes the person the organisation associates with CS and AI.

But sharing changes the stakes completely. When your personal twin gets something slightly wrong, you catch it, you know its quirks. When a shared assistant gets something wrong, it tells the whole team confidently, and people who do not know its quirks act on it. The blast radius of an error goes from one to everyone. So a shared knowledge base is not just a bigger personal one. It is a different thing, with a different bar for accuracy.

The point

A shared knowledge base turns your fluency into company capability and you into the authority. But its errors reach everyone confidently, so one wrong shared answer is worse than no answer. The bar for accuracy is higher than anything you build for yourself alone.

2
Designing a knowledge base that stays right

A knowledge base is only trustworthy if it is accurate, and accuracy is not a one-time state, it is something that decays as products change, prices change, and playbooks evolve. The design job is building one that stays right, not one that is right on launch day and rots after.

Three design choices do most of the work. Single source of truth: each fact lives in exactly one place, so when it changes you update it once, not in five documents that then disagree. Clear ownership of each area: someone is responsible for keeping the pricing section, the product section, the process section current, with you owning the whole. And a freshness signal on everything: a last-reviewed date, so anyone (and the assistant) can see what is current and what might be stale. A knowledge base without these is a confident liar in waiting, accurate today and quietly wrong in three months.

The design rule

Build for staying right, not being right on day one. One source of truth per fact, clear ownership of each area, and a last-reviewed date on everything. Accuracy decays. The design has to fight that decay, or the base becomes a confident source of wrong answers.

3
Owning the curation is what makes you the authority

Here is the part people miss: the authority does not come from building the knowledge base. It comes from owning its accuracy over time. Anyone can stand up an assistant in an afternoon. The person who becomes indispensable is the one who keeps it trustworthy, who reviews the stale sections, who catches the wrong answer before it spreads, who is responsible when someone asks 'can I trust what this thing told me?'.

That ownership is a real job, not a side effect, and naming it as your job is how you claim the authority. It also protects the team: a shared assistant with no owner is the dangerous version, because no one is watching for the wrong answer that reaches everyone. By taking the curation role, you make the capability safe and make yourself the person the organisation depends on for it. The two go together, the value to them and the value to you are the same act.

What you walk away with

A shared knowledge base design plus a working team-assistant prototype, and the curation role that makes you the authority. The building is the easy part. Owning the accuracy over time is what makes you indispensable and keeps the team safe.

What happened

A team CS assistant, built helpfully by someone, answers a rep's question about a product's price using a figure from a document loaded six months ago. The price changed three months ago. The rep, trusting the assistant, quotes the old price to a customer. Three other reps had quietly done the same that month. The error did not happen once, it happened to the whole team, confidently, for months.

What broke, and whose job it was

Two things broke. First, design: the old price lived in a loaded document with no last-reviewed date and no single source of truth, so when the price changed, the stale figure stayed in the knowledge base contradicting reality, and nothing flagged it as old. Second, and the deeper failure: no one owned the accuracy. The assistant had a builder but not a curator. So there was no one whose job it was to update the price when it changed or to catch that the base had gone stale. The wrong answer reached everyone precisely because no one was responsible for keeping it right.

The fix, and the lesson

The design fix is structural: one source of truth for each price, a last-reviewed date on every fact, and a regular review of the sections most likely to change. But the real fix is ownership. Someone, the CS AI authority, takes responsibility for the base staying accurate, including a process for updating facts when they change and reviewing stale areas. That ownership is exactly the role that makes that person the authority. The lesson cuts both ways: a shared base without an owner is a liability, and being the owner is how you turn the capability into your standing.

Practice check · not scored
Try the judgement call
A teammate proudly builds a shared CS assistant in an afternoon by uploading a big folder of current documents. It works well in the demo. You are asked whether it is ready for the whole team to rely on.

What is the most important thing to raise before it goes live?

The building is the easy part, the demo proves the happy path. The danger of a shared assistant is that its errors reach everyone confidently, and accuracy decays as facts change, so the critical question is ownership: who keeps it right and how stale facts get caught. Without that, it becomes a confident liar over time. 'Works in the demo' ignores decay, model power does not prevent stale facts, and polish is cosmetic next to accuracy.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why is a shared team knowledge base held to a higher accuracy bar than your personal twin?

The blast radius is the key idea. You know your personal twin's quirks and catch its slips, but a shared assistant tells the whole team confidently, and colleagues who do not know its quirks act on the answer, so one wrong shared answer is worse than none. It is not about different technology, the difference is real, and the higher bar is on the shared base precisely because of who relies on it and how.

2. Which set of design choices best keeps a shared knowledge base accurate as products and prices change over time?

Accuracy decays, so the design must fight decay: a single source of truth (update once, no contradicting copies), clear ownership (someone keeps each area current), and freshness signals (last-reviewed dates) that reveal what is stale. Loading more documents increases contradictions and stale copies, a newer model cannot reason its way past a wrong loaded fact, and locking the base guarantees it goes stale as reality moves on.

3. What actually makes someone the recognised CS AI authority in their company?

Anyone can build an assistant in an afternoon, so building alone does not create authority. Authority comes from owning the accuracy over time, being the person responsible for the capability staying trustworthy, which is also what keeps the team safe. Being first or most complex does not make you depended upon, and keeping methods private (the opposite of the next module's lesson) builds a hoarder's fragile edge, not genuine authority.

Score: 0/3 ·

You will leave able to

  • Design a knowledge base that stays right even as facts change
  • Build a shared assistant the team can genuinely trust
  • Own the curation role that makes you the company's CS AI authority

Hands-on exercise

Design a small shared CS knowledge base: pick three areas (say product facts, pricing, a key process), give each a single source of truth and an owner, and put a last-reviewed date on each. Build a prototype team assistant on it. Then write down the curation process: who updates what, how often stale areas get reviewed, and who is responsible when the answer is wrong. That last part is the authority.

The human element: Owning a shared capability is owning a responsibility to your colleagues, not just claiming a title. The authority is real precisely because the duty is real: people are trusting what your system tells them. Take the title only if you will do the curation that earns it.

If you remember three things

  1. A shared knowledge base turns your fluency into company capability, but its errors reach everyone, so the accuracy bar is higher.
  2. Design for staying right: one source of truth per fact, clear ownership, and a last-reviewed date on everything.
  3. Authority comes from owning the accuracy over time, not from building it. That ownership keeps the team safe and makes you indispensable.
Part 3 · Lead · Module 11

Keep what you built working, and defend it

Testing, catching silent decay, and defending any output when someone pushes back

14 min 85% to pass

The lesson

30days of silent decay

Anyone can build a clever thing once. The real skill is keeping it working: testing your systems properly, noticing when something has quietly got worse after a model update, and being able to defend any answer when an executive or auditor pushes back. This is what makes your systems safe to lean on.

1
Built once is not built well

There is a gap between someone who made a clever thing and someone who runs a reliable system, and the gap is maintenance. The systems you have built, the twin, the agent, the orchestrated routine, the shared knowledge base, all decay the same way prompts do: models update, accounts change, facts go stale. A system that was excellent at launch and unmaintained for six months is not a slightly-worse system, it is an unknown one, because nobody knows whether it still works.

This matters more for systems than for single prompts, because systems act and systems are trusted by others. A decayed prompt gives you one weaker answer. A decayed agent quietly does the wrong thing for weeks, and a decayed shared base tells the team wrong answers confidently. The whole value of what you built rests on it staying reliable, and reliability is a practice, not a launch state.

The point

The difference between a builder and someone who made a clever thing once is maintenance. Your systems decay silently as models and facts change, and because they act and are trusted, decay is more dangerous than for a single prompt. Reliability is a practice you keep, not a state you reach.

2
Catching silent decay before it costs you

Silent decay is the dangerous kind, because nothing announces it. The system keeps running, keeps producing confident output, and just gets quietly worse. The only defence is the same one from your prompt library, scaled up: a test, run regularly, that would reveal decay.

For each important system, keep a known-good test: a saved input and a record of what good output looked like. Re-run it on a schedule, monthly, say, and after any model update, and compare. If the output has drifted from the known-good example, you have caught decay before it cost you, and you can find and fix the cause. The thing that makes this work is having captured 'good' when the system was working, you cannot recognise decay if you never recorded the baseline. The test is cheap. The thirty days of a decayed agent running unnoticed, the example this module opened with, is not.

The discipline

Decay is silent, so you cannot wait to notice it, you have to test for it. Keep a known-good baseline for each system and re-run it on a schedule and after model updates. You can only recognise decay against a baseline you recorded when things worked.

3
Defending what you built

The more your systems matter, the more someone will eventually challenge them: an executive asking 'how do you know this is right?', a security or governance review asking 'what does this thing touch and how is it controlled?', a colleague asking 'can I trust what it told me?'. Being able to answer is part of owning the system, not an afterthought.

Defensibility is built, not improvised in the meeting. For any system that matters you should be able to say, on one page: what it does, what data it touches and where that data lives, what it must never do, where a human checks it, and how you know it still works (your test). A system you can explain that clearly is one you can defend, and one others can trust. A system you cannot explain that clearly is a risk you have not yet understood, however well it runs. The act of writing the one-page defence often reveals a gap you then fix, which is the point: defensibility and safety are the same thing seen from two angles.

What you walk away with

A test-and-maintenance plan for everything you built, plus a one-page defence of your main system: what it does, what it touches, what it must never do, where the human checks are, and how you know it still works. If you can write that page, you can defend it. If you cannot, you have found the gap to fix.

What happened

Your portfolio agent ran every Monday and was trusted by you and two colleagues you had shared it with. A model update landed mid-quarter. The agent kept running and kept producing confident, polished watchlists. But its risk reasoning had subtly weakened, it stopped reliably doing the disconfirming check, so it under-flagged a couple of genuinely at-risk accounts. Nobody noticed for a month, because the output still looked exactly as confident as before.

Why it went unnoticed, and what would have caught it

It went unnoticed because nothing was testing for decay, the system was trusted and left to run, and decayed output looks identical to good output on the surface. What would have caught it: a known-good test, a saved real input with a record of the strong analysis it used to produce, re-run after the model update. Comparing the new output to the recorded baseline would have immediately shown the disconfirming step had dropped out, catching the decay in days instead of a month.

The fix, and defending it afterwards

The immediate fix is the module-5 move: re-state the disconfirming check as a hard, non-optional instruction so the model update cannot quietly drop it, then re-run the test to confirm the baseline is restored. The lasting fix is the practice: that test now runs after every model update. And when your manager asks 'how do we know your agent is reliable?', you can answer with the one-page defence, what it does, what it touches, where you check it, and the test that proves it still works, which is exactly the answer that turns a nervous question into trust. The decay was a problem. Being able to show how you catch and fix it is what makes you defensible.

Practice check · not scored
Try the judgement call
A system you built and shared has been running confidently for weeks. A colleague asks how you know it has not quietly degraded since the last model update. You realise you have no way to answer.

What was the missing practice, and what fixes it going forward?

Decay is silent and decayed output looks as confident as good output, so 'runs without errors' proves nothing about quality. The missing practice is a known-good baseline, a saved input and a record of good output, re-run on a schedule and after updates, so you can detect drift against a recorded standard. Model updates do not only improve things (they can drop behaviours), and rebuilding from scratch each time is wasteful when a test plus a targeted fix does the job.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why is silent decay more dangerous in a built system (an agent or shared base) than in a single saved prompt?

The danger scales with action and trust. A decayed prompt gives you one weaker answer that you, the user, can judge. A decayed agent acts wrongly for weeks, and a decayed shared base feeds the whole team confident wrong answers, all silently. It is not about different technology, prompts can and do decay too (that was module 2), and the difference in danger is real and comes from systems acting and being relied on by others.

2. What single thing makes it possible to detect that a system has decayed?

You can only recognise decay against a baseline you recorded when things worked, which is why capturing 'good' output up front is essential. Decayed systems usually run without errors and still look confident and polished, that surface normality is exactly why decay is silent, so neither error-watching nor 'looks confident' detects it. And asking the model to self-report its degradation is not reliable, the comparison to a recorded baseline is.

3. What is the practical test of whether a system you built is defensible to an executive or auditor?

Defensibility is the ability to clearly account for the system: its purpose, the data it touches and where that lives, its hard limits, its human checkpoints, and the test proving it still works. If you can write that one page, you can defend it and others can trust it, if you cannot, you have found a gap. Model choice, an error-free run so far, and speed say nothing about whether you can actually account for what the system does and touches.

Score: 0/3 ·

You will leave able to

  • Build a repeatable test that would reveal a system has decayed
  • Catch silent degradation after a model update before it costs you
  • Defend any system you built to a sceptical executive or auditor

Hands-on exercise

For your main built system, capture a known-good baseline now: a saved real input and a record of the good output it currently produces. Schedule a re-run, monthly and after any model update. Then write its one-page defence: what it does, what data it touches and where, what it must never do, where you check it, and how the test proves it works. Note any gap the writing reveals, and fix it.

The human element: Defending your work is not bureaucracy, it is the proof that you understand what you built. The systems you can explain clearly are the ones you truly own. The ones you cannot are quietly owning a piece of your risk, however smoothly they run today.

If you remember three things

  1. Maintenance is what separates a builder from someone who made a clever thing once. Systems decay silently as models and facts change.
  2. You cannot wait to notice decay, you have to test for it against a known-good baseline captured when things worked.
  3. Defensibility and safety are the same thing: if you can explain on one page what a system does, touches, refuses, checks, and how you know it works, you can defend it and others can trust it.
Part 3 · Lead · Module 12

Turn what you built into real standing

Sharing, teaching, and positioning so your work becomes a reputation, not a private trick

14 min 85% to pass

The lesson

giveit away to gain

The last step is turning all this into a reputation. How to share what you have built so it spreads and you still get the credit, how to teach others, how to have a say in your company's AI direction, and how to build a name as a CS AI authority, inside your company and in the wider field.

1
Capability is not the same as standing

You now run a measurably different operation from most of your peers. But here is the trap that catches good builders: capability does not automatically become standing. The CSM who quietly has brilliant systems and tells no one is more productive, and exactly as replaceable as before in the eyes of the organisation, because nobody knows. Skill kept private is a personal convenience. Skill made visible and useful to others is authority.

This module is the deliberate work of converting what you can do into what you are known for. It is not bragging, and it is not optional if you want the career return on all this effort. The systems you built are the input. Standing is the output, and the conversion does not happen by itself. You have to do it on purpose.

The point

Capability kept private is a convenience. Capability made visible and useful to others is authority. The conversion from one to the other does not happen by itself, and skipping it means doing all the work and getting none of the standing.

2
Share and teach, the counterintuitive move

The instinct is to guard your edge: if you built something great, keeping it secret protects your advantage. This is exactly backwards. The hoarder is a CSM with a private trick, useful until someone else builds the same thing, and then ordinary again. The teacher becomes the person the whole organisation associates with the capability, which is durable in a way a private trick never is.

So give it away on purpose. Run the lunch-and-learn. Share your library. Walk a colleague through building their first agent. Quantify the impact in a way leadership remembers, not 'I use AI a lot' but 'I recovered roughly six hours a week and reinvested them in my top accounts, here are the two saves and the expansion that came from it'. Recovered hours are the input, the saves and growth are the story. When every CS team is being asked 'what is our AI plan?', the person who has been visibly teaching the answer is worth more than any private edge could ever be.

The counterintuitive truth

Giving the methods away beats hoarding them. The hoarder has a trick that goes ordinary the moment someone copies it. The teacher becomes the person the organisation depends on for the whole capability, which only grows as more people use it.

3
Positioning, inside and out

Two arenas, and both compound. Inside your company: be the voice in the room when AI direction is discussed, the one who has actually built things and can speak from real experience rather than slideware. Volunteer to shape the team's approach, the standards, the rollout. The person who built and teaches the capability is the natural person to help steer it, and steering it is a leadership role whether or not it has a title yet.

Outside your company: the wider CS field is hungry for people who have genuinely done this, not just talked about it. A post about a real thing you built and what you learned, a talk at a meet-up, a written breakdown of a system, these build a professional profile that outlasts any single job. The same act, done in public, turns internal standing into a reputation in the field. You have spent this whole course building real things. This final step makes sure that the work becomes who you are known as, which is the truest form of career capital there is.

What you walk away with

A plan to roll one of your builds out to your team, and a plan to position yourself, internally as the voice on AI direction, externally through sharing real work. The systems were the hard part. This is making sure the work becomes your reputation.

The hoarder

One CSM builds excellent systems and keeps them to themselves, worried that sharing would erode their edge. For a while they are quietly the most productive person on the team. Then the company rolls out AI training, three colleagues build similar systems, and the edge evaporates. A year on, they are a good CSM with good tools, and no one particularly associates them with AI. The capability was real, the standing never materialised, because nobody knew.

The teacher

Another CSM builds the same kind of systems and does the opposite: runs a lunch-and-learn, shares the library, helps two colleagues build their first agents, and tells the value story to leadership with real numbers. When the company asks 'what is our CS AI plan?', leadership turns to them, because they are visibly the person who has been doing it. They get asked to shape the team's approach. A year on, they have a title conversation, a profile in the wider CS community from a couple of posts about real builds, and a reputation that would follow them to any job.

The lesson

Both built equally good systems. The only difference was what they did with the visibility, and it produced completely different careers. The hoarder optimised for a private edge and got a temporary one. The teacher optimised for being known for the capability and got durable authority. Giving it away did not cost the teacher their advantage, it converted a fragile private edge into a standing that grows as more people rely on them. The work was necessary. Making the work visible and useful to others is what turned it into a career.

Practice check · not scored
Try the judgement call
You have built genuinely strong AI systems over this course. A colleague advises you to keep them private so you stay ahead of the team and protect your advantage.

Why is that advice likely to cost you in the long run?

A private trick is fragile: it lasts only until someone else figures out the same thing, then you are ordinary again with no standing to show for the work. Teaching and sharing make you the person the organisation associates with the whole capability, which is durable and grows with adoption. It is not about a policy requirement or any doubt about your systems, it is that visibility converts capability into authority while secrecy leaves you replaceable.

You've reached the knowledge check. Scenario-based questions. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. What is the core reason that capability does not automatically translate into standing?

The conversion gap is visibility. You can run brilliant systems and remain exactly as replaceable in the organisation's eyes as before, because nobody knows, so the skill stays a private convenience rather than becoming authority. Standing does not require a title to begin (you can be the de facto voice on AI first), the two are not the same thing (that is the whole point), and it is about visibility and usefulness to others, not extra technical skill.

2. Why does teaching and sharing your methods build more durable authority than keeping them as a private edge?

Durability is the point. A private edge is temporary, it ends the moment a colleague or a company rollout reproduces it. Being the person who teaches and is associated with the capability is durable and compounds: the more people use what you spread, the more central you become. Teaching is not less effort, sharing does not stop others learning (it spreads learning), and private methods are not inherently inferior, they are just fragile as a source of standing.

3. When telling leadership the value of your AI work, which framing builds the most standing?

Leadership remembers value tied to outcomes they care about: time recovered and reinvested into saves and growth. Recovered hours are the input and the saves and expansion are the story, which is the framing that lands. 'I use AI a lot' and 'I built the most systems' are activity claims, not value, and a technical description of how the systems work is for builders, not for leadership deciding what your work is worth.

Score: 0/3 ·

You will leave able to

  • Share a build so it spreads and you are still credited for it
  • Teach the capability in a way that lands and tell the value story leadership remembers
  • Position yourself as a recognised CS AI authority, inside your company and in the wider field

Hands-on exercise

Pick one build you are proud of. Plan to give it away: a short lunch-and-learn or a shared, documented version a colleague could adopt. Write your value story in leadership's language, recovered hours reinvested into specific saves and growth. Then draft one public post or talk outline about a real thing you built and what you learned. That is capability becoming reputation.

The human element: Authority built on giving real value to others is the kind that lasts and the kind worth having. The hoarded edge buys a year. Being genuinely, visibly the person who helped everyone get better buys a career, and it is the more generous way to live in a team besides.

If you remember three things

  1. Capability kept private is a convenience. Made visible and useful to others, it becomes authority. The conversion does not happen by itself.
  2. Giving your methods away beats hoarding them: a private edge goes ordinary, a taught capability makes you durably central.
  3. Position in both arenas: be the internal voice on AI direction, and build an external profile by sharing real work. The work becomes your reputation.
Advanced Certification

Built, not attended.

This certificate says you can build AI systems for Customer Success and exercise expert judgement about when to trust, overrule, or keep AI out. Earned by passing twelve hard knowledge checks at 85 percent.

1

Work through the 12 modules

Build yourself, build systems, lead. Each module ends in a hands-on build and a scenario knowledge check.

2

Pass every knowledge check at 85%

Harder than the foundational course. Wrong answers are explained, not just marked. A module counts as complete only once its check is passed.

3

Generate your certificate

Finish all 12 and the certificate unlocks: enter your name and download it, dated and stamped with a unique ID. Made for your LinkedIn profile.

Sample

This is the actual certificate, rendered live. Yours downloads in full resolution as a landscape image and a square version made for LinkedIn.

Complete all 12 modules to unlock your certificate.

For team leads and CS managers

Take your team from scattered AI use to a real capability.

Nine modules on the actual job of leading a team into AI: the trust gap, the data line, the rollout that does not stall, proving value, and protecting the human core. Grounded in what CS leaders are really facing, not theory.

Capability
Trust
0/9

Modules complete

No account, no tracking, nothing leaves your machine. Your progress lives in this browser, so stick to the same browser to keep your place. The knowledge checks pass at 85 percent.

Part 1 · Get honest

See your team and yourself clearly

M01–M03 · your real job, the shadow use, the trust gap

MODULE 01

Your real job is not using AI, it is leading it

Why leading a team through AI is a different job from being good at it yourself

MODULE 02

Read what is really going on, including the shadow use

Seeing the true state of your team's AI use, not the official version

MODULE 03

The trust gap, and why your team is more nervous than you think

Closing the gap between how much you trust AI and how much your team does

Part 2 · Build capability

Move the whole team forward

M04–M06 · the data line, the rollout, consistent quality

MODULE 04

Draw the line you can actually defend

A one-page data and tool boundary your team can follow and you can defend upward

MODULE 05

Get the whole team using it, not just the keen few

The rollout sequence that moves past the two enthusiasts to real, consistent use

MODULE 06

Get consistent quality without crushing your people

Raising the floor without lowering the ceiling, and coaching the over-reliant

Part 3 · Prove and protect

Hold the trust upward and the humanity inward

M07–M09 · prove the value, governance, lead through change

MODULE 07

Prove the value in language your leadership respects

Measuring what matters and reporting it honestly, so you keep the trust and the investment

MODULE 08

Be ready for the governance question before it is asked

Accountability and oversight so an AI-assisted decision can always be explained

MODULE 09

Lead the change without losing the human core of the job

Protecting what must stay human, and leading people through real anxiety honestly

Part 1 · Get honest · Module 01

Your real job is not using AI, it is leading it

Why leading a team through AI is a different job from being good at it yourself

14 min 85% to pass

The lesson

2different jobs

Being personally good with AI and leading a team through it are two different jobs. Plenty of skilled individual users make poor AI leaders, because the leader's job is making the good way the easy, default way for everyone, and owning the risk when it goes wrong.

1
The skill that does not transfer

Here is the trap that catches good CS leaders: they assume that because they are personally fluent with AI, they can lead their team into it. The two are barely related. Being good at AI is about your own output. Leading AI is about everyone else's: their habits, their fears, their shortcuts, and the risk they carry into your accounts. A brilliant individual user can run a chaotic, exposed team, and a modest individual user can run a calm, capable one. The skills are different.

Your job as the leader is not to be the best prompter in the room. It is to make the good way, the safe way, the consistent way, the easy and default way for everyone, so that the right behaviour happens without anyone having to be a hero about it. And it is to own the risk when something goes wrong, because it will. That is leadership work, not power-user work.

The point

Your personal AI skill and your job of leading AI are different things. The leader's job is making the good way the default way for the whole team, and owning the risk. Being the best prompter in the room is not the job, and mistaking it for the job is how capable leaders run exposed teams.

2
The three things only you can own

Most of the work of AI adoption can be shared, taught, or delegated to an enthusiast on the team. Three things cannot. These are the parts that are yours, and naming them is the start of doing the job well.

What only the leader can own
leadership rests on these three 1 The line 2 The tone 3 The risk

Everything else can be delegated. Take any one of these away and the team is exposed.

🧭
The line

What data can go into which tools, and where the boundary is. Only you can set a line the team can actually follow and that you can defend upward. An enthusiast cannot own this, because it carries real risk and requires authority.

🌡️
The tone

Whether the team feels safe to be honest about how they really use AI, or hides it. You set that quietly, by how you react the first time someone admits to a shortcut or a mistake. The tone decides whether you ever see the truth.

🛡️
The risk

When something goes wrong, and it will, you own it, not the CSM who pasted the wrong thing or the tool that invented a figure. Owning the risk is what lets you set the line and the tone with authority. It is the price of leading.

What only you can do

The line, the tone, and the risk. Everything else, training, prompts, even running sessions, can be shared or delegated. These three are yours, because they require authority and they carry consequences. If you do not own them, no one does, and the team is exposed.

3
Tone is the quiet decider

Of the three, tone is the one leaders underrate most, because it is invisible and never appears in a plan. But it is what decides whether your adoption works or quietly fails. If your team believes that admitting to using an unapproved tool, or to leaning too hard on AI, will get them in trouble, they will simply stop telling you. And the moment you cannot see what is really happening, you cannot lead it. You are managing a fiction.

The tone is set in small moments, not announcements. The first time someone admits they have been using a tool you did not approve, your reaction teaches the whole team whether honesty is safe. React with curiosity ("interesting, why is that one better?") and you keep visibility. React with punishment and you lose it, permanently, and the shadow use does not stop, it just goes quiet. Setting a tone where the truth surfaces is not being soft. It is the only way to actually lead what is really happening rather than the official version of it.

What you walk away with

A clear picture of your actual role, the three things only you can own (the line, the tone, the risk), and the understanding that tone is the quiet decider. The leader who makes honesty safe sees the truth and can lead it. The leader who punishes it manages a fiction.

Team A: thriving

Same budget, same approved tools as Team B. Six months in, Team A uses AI consistently and safely, the leader can say exactly what the team uses it for and where the risks are, and quality is steady. The leader is not the best individual user on the team, two of their CSMs are sharper with prompts. But the leader did three things: set a clear, plain data line everyone understood, made it visibly safe to be honest about real usage, and took personal responsibility the one time something went wrong rather than hunting for who to blame.

Team B: chaos

Same tools, same budget. Six months in, the leader genuinely does not know what their team uses AI for. Some are using unapproved tools quietly. Quality swings wildly. There was a near-miss with customer data that the leader only found out about by accident. This leader is personally excellent with AI, better than anyone on Team A. But they treated leading it as a technical problem, focused on showing off good prompts, never set a clear line, and the one time someone admitted a mistake, they made an example of them, after which the team went silent.

What the first leader did that the second did not

The difference was not skill, talent, tools, or budget, those were equal or favoured Team B. The difference was that Team A's leader did the actual leadership job: owned the line, set a tone where the truth surfaced, and carried the risk personally. Team B's leader did the power-user job, being impressive with AI, while leaving the leadership work undone. The lesson is the whole module: leading AI is not being good at AI, and confusing the two produces a brilliant individual presiding over an exposed, silent team.

Practice check · not scored
Try the judgement call
A CS leader is the most AI-skilled person in their company and runs impressive demos for the team. Six months in, they cannot actually say what their team uses AI for, two people are quietly using unapproved tools, and a data near-miss surfaced by accident.

What is the core problem?

Personal skill was not the gap, they were the most skilled person in the company. The gap is that being good at AI and leading a team through it are different jobs, and they did only the first. The leadership job, setting a defensible line, making honesty safe so they can see real usage, and owning the risk, was left undone, which is exactly why they cannot see what their team does. Better personal skill, a 'better' team, or different tools would not fix a leadership gap.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why can a CS leader who is personally excellent with AI still end up running an exposed, chaotic team?

The two jobs are barely related: personal skill is about your own output, while leading AI is about everyone else's habits, fears, shortcuts, and the risk they carry. A brilliant individual user who never sets a line, a safe tone, or owns the risk will run an exposed team. It is not that skill transfers, that excellent users lack time, or that the team is jealous, it is that the leadership work is a separate job that was left undone.

2. Of the line, the tone, and the risk, why is tone the one leaders most often underrate?

Tone is underrated precisely because it is invisible and never shows up in a rollout plan, yet it determines whether people are honest about their real usage or hide it, and you cannot lead what you cannot see. It matters enormously, it is set in small moments not announcements, and it applies to leaders at every level. Underrating it is how leaders end up managing the official version of reality instead of the real one.

3. A team member admits they have been using an AI tool you never approved. What does your reaction in that moment actually determine?

These small moments set the tone for everyone, not just the individual. React with curiosity and you keep visibility into real usage, react with punishment and the whole team learns to go quiet, so the shadow use continues but invisibly. It is far from a one-off with no consequence, and jumping to removing the person is exactly the punishing reaction that destroys honesty and your ability to lead what is really happening.

Score: 0/3 ·

You will leave able to

  • Separate your personal AI skill from your job of leading it
  • Name the three things only the leader can own: the line, the tone, the risk
  • Set the tone that quietly decides whether adoption works or stalls

Hands-on exercise

Write down, honestly, how much of your AI energy goes into your own output versus leading the team's. Then write your answers to the three things only you can own: where is your data line, what tone have you actually set about honesty (judged by how people behave, not what you have said), and do you truly own the risk or quietly hope it lands elsewhere. The gaps are your real job.

The human element: Leading people through change is older than AI and has not changed: it runs on trust, honesty, and someone willing to carry the risk. The tool is new, the leadership is not. The leaders who do this well are the ones who remember that.

If you remember three things

  1. Being good at AI and leading a team through it are different jobs. Confusing them produces a brilliant individual running an exposed team.
  2. Three things only you can own: the line (what data goes where), the tone (whether honesty is safe), and the risk (you carry it).
  3. Tone is the quiet decider. Make honesty safe and you see the truth and can lead it. Punish it and you manage a fiction.
Part 1 · Get honest · Module 02

Read what is really going on, including the shadow use

Seeing the true state of your team's AI use, not the official version

14 min 85% to pass

The lesson

78% use unapproved tools

Right now, some of your team are using AI tools nobody approved, and they are not telling you. Surveys put unsanctioned use at most teams. That is not a sign they are bad people, it is a sign the approved options are not good enough, so they found their own. You cannot lead what you cannot see.

1
The official version is not the real one

Every CS leader has two pictures of their team's AI use: the official one (what is approved, what people say in meetings) and the real one (what people actually do at 4pm on a deadline). The gap between them is where your risk and your opportunity both live, and most leaders never see the real picture because nothing in their normal reporting reveals it.

The single most important fact to absorb: a large majority of employees use AI tools their organisation has not approved, and they do not advertise it. This is not a few rogue actors, it is the normal state of a modern team. So if your picture of your team's AI use is the official one, it is wrong, not because people are lying to you exactly, but because nobody volunteers the shortcut they took under pressure. Leading the official version while the real version runs underneath is the most common failure in this whole course.

The point

You have two pictures of your team's AI use: the official one and the real one. Most leaders only ever see the official one, which means they are leading a fiction. Seeing the real picture, shadow use and all, is the precondition for leading any of it.

2
Shadow use is a signal, not just a violation

When you discover someone is using a tool you never approved, the instinct is to treat it as a rule broken. That instinct, acted on, destroys your visibility (last module's tone lesson) and misses the more useful reading. Shadow use is information. Someone went out of their way to find and use an unapproved tool, which almost always means the approved option you gave them is not good enough for the job. The rule-breaking is the symptom. The inadequate approved tool is the cause.

So the first question on discovering shadow use is not "how do I stop this?" but "what is this telling me about the gap between what my team needs and what I have given them?" Sometimes the answer is that you need a better approved tool. Sometimes it is that the approved tool is fine and people did not know how to use it. Either way, the shadow use has handed you a genuine piece of intelligence about your own provision, for free, and a leader who only sees a rule broken throws that intelligence away along with their visibility.

The reframe

Shadow use is a signal about your approved options, not just a rule broken. Someone going out of their way to use an unapproved tool usually means what you gave them is not good enough. Read the cause, not just the symptom, or you throw away real intelligence about your own provision.

3
Reading the patterns on your team

Once you can see the real picture, you can read the patterns, and different patterns need different responses. The four to learn to tell apart are below. The most dangerous mistake is treating them all the same, because the person racing ahead and the person quietly drowning can produce similar surface numbers.

🚀
Racing ahead

Using AI heavily and capably, often including shadow tools. An asset and a risk: they are your potential champions, but they may be carrying data risk or skipping judgement. Harness them, do not just celebrate or punish them.

😈
Quietly at risk

Also using AI heavily, but leaning on it past their judgement, output looks fine, underlying skill fading. Looks like 'racing ahead' on the numbers. The hardest to spot and the one that costs you a renewal later.

🥶
Frozen

Barely using it, not from defiance but from not knowing how or not trusting it. Needs support and a safe on-ramp, not pressure. Pressure turns frozen into resentful.

😨
Scared

Avoiding it because they fear it threatens their job. Looks like 'frozen' but the cause is different and the fix is the trust work in the next module, not a tutorial.

What you walk away with

An honest map of your team's real AI use, including the shadow use, and the ability to read the four patterns apart. The two that look alike, racing ahead versus quietly at risk, and frozen versus scared, need opposite responses. Treating them the same is the costly mistake.

What you discover

You find out, almost by accident, that around half your team has been quietly using a particular AI tool you never approved, because it is genuinely better at a common task than the approved one you gave them. Your first instinct is that a rule has been broken across half the team and you need to clamp down.

What it is actually telling you

Clamping down would be the costly misread. Half your team independently going to the same unapproved tool is not a discipline problem, it is the loudest possible signal that your approved tool is not good enough for that task. They are not being reckless for fun, they found something that helps them do their job and used it. The shadow use has just handed you precise intelligence: here is exactly where your provision is failing your team. That is valuable, and a leader who only sees the rule broken throws it away.

What a good leader does first

Not punish, and not ignore either, both are wrong. First, understand: talk to people with genuine curiosity about why that tool is better, treating it as intelligence, which also protects the honesty you need. Then act on the real problem in two tracks at once: address the data risk (the unapproved tool may be exposing customer data, which is the genuine concern, not the rule itself) by being honest about why the line exists, and move to close the provision gap, either get the better tool approved or fix how people use the approved one. The rule-breaking is handled by fixing its cause, not by punishing the symptom and driving it underground.

Practice check · not scored
Try the judgement call
You discover that several of your strongest CSMs have quietly adopted an unapproved AI tool because it genuinely outperforms the approved one for a key task.

What should your first move be?

Strong people independently choosing an unapproved tool is a signal that your approved provision has a gap, so the first move is to understand it, which also preserves the honesty you depend on. You then act on the real concerns: the data risk (the genuine reason for a line) and closing the provision gap. A warning punishes the symptom and drives it underground, ignoring it leaves real data risk unaddressed, and approving the tool without checking its data handling skips the actual safety question.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why is leading from the 'official' picture of your team's AI use a serious mistake?

A large majority of employees use unapproved tools without advertising it, so the official picture systematically understates and misrepresents real use. Leading that fiction means missing both your risk and your opportunity. It is not that the official picture overstates use or is perfectly accurate, and the answer is not to stop tracking, it is to get sight of the real picture, shadow use included.

2. What is the most useful way to read the discovery that someone is using an unapproved tool?

Shadow use is intelligence: someone going out of their way to find an unapproved tool usually means your approved provision has a gap, so the cause is an inadequate tool, not a character flaw. Reading it as untrustworthiness destroys your visibility and misses the signal, dropping data rules entirely is an overcorrection that ignores real risk, and silently logging it addresses neither the cause nor the risk.

3. Two CSMs both show heavy AI usage in your numbers. Why might they need opposite responses from you?

'Racing ahead' and 'quietly at risk' can produce near-identical surface numbers, but one is an asset to harness and the other is a person whose judgement is eroding and will cost you a renewal, so they need opposite responses. Identical usage does not mean identical situations, seniority is not the test, and telling both to use AI less wrongly treats the capable one as a problem while missing the real issue with the at-risk one.

Score: 0/3 ·

You will leave able to

  • See the real state of your team, not the official version
  • Read unapproved tool use as a signal about your provision, not just a rule broken
  • Tell the difference between someone racing ahead and someone quietly at risk

Hands-on exercise

For each person on your team, write down which of the four patterns they are in: racing ahead, quietly at risk, frozen, or scared. Be honest about how confident you actually are, and notice where you genuinely do not know, that uncertainty is the gap in your visibility. Then, without punishing anyone, find one honest way to learn what tools your team is really using.

The human element: People hide their shortcuts from leaders they do not feel safe with, and show them to leaders they trust. The amount of shadow use you can see is a direct measure of how safe your team feels with you. Low visibility is not a people problem, it is feedback.

If you remember three things

  1. You have an official picture and a real picture of your team's AI use. Most leaders only see the official one and so lead a fiction.
  2. Shadow use is a signal that your approved tools have a gap, not just a rule broken. Read the cause, not the symptom.
  3. Learn to tell the four patterns apart. Racing ahead versus quietly at risk, and frozen versus scared, look alike but need opposite responses.
Part 1 · Get honest · Module 03

The trust gap, and why your team is more nervous than you think

Closing the gap between how much you trust AI and how much your team does

16 min 85% to pass

The lesson

9% vs 61%

An uncomfortable truth from the research: leaders trust AI for important decisions far more than their frontline teams do, by a wide margin. So when you push adoption, you are often asking people to lean on something they quietly do not trust, while some also fear it is coming for their jobs.

1
The gap you cannot see from where you sit

Here is a number worth sitting with: surveys consistently find that executives trust AI for high-stakes decisions far more than frontline staff do, a gap of fifty points or more in some studies. You, as the leader, are much closer to the executive end of that gap than your team is. So your own comfort with AI is actively misleading you about how your team feels, you are standing on the confident side of a divide and assuming everyone is there with you.

The gap you are standing across
Your team
9%

trust AI for important decisions

You, the leader
61%

trust AI for important decisions

50+ point gap you cannot see from your seat

You experience pushing adoption as encouraging an obviously good thing. Your team experiences it as being told to lean on something they do not trust.

This matters because of what it does to adoption. When you push your team to use AI more, you experience it as encouraging an obviously good thing. A nervous team member experiences it as being told to lean on something they do not trust, possibly by someone who does not understand why they are hesitant. The gap is invisible from your seat, which is exactly why rollouts stall in ways that baffle the leader: you are solving a tools problem while your team has a trust problem.

The point

Leaders trust AI far more than their frontline teams do, and you are on the confident side of that gap. Your own comfort misleads you about how your team feels. Push adoption without closing the trust gap and the rollout stalls in ways you will not understand, because you are solving the wrong problem.

2
The job-fear question, answered honestly

Underneath a lot of the hesitation is a question people rarely say out loud: is this coming for my job? If you do not address it, it sits there and quietly poisons every adoption effort, because no one leans into the tool they believe is replacing them. And the way most leaders handle it makes it worse: they reach for the comfortable platitude, "don't worry, AI won't replace you," which everyone correctly hears as either naive or dishonest.

The honest answer is more useful and more respected, even though it is harder to give. It is some version of: the work is genuinely changing, the parts that are routine will increasingly be done with AI, and that is real, not something I will pretend away. What stays valuable, and becomes more valuable, is the judgement, the relationships, and the accountability that AI cannot provide. My job as your leader is to help you move up into that work, not to pretend nothing is happening. That answer treats people as adults, names the real change, and points to a real future. The platitude treats them as children and earns no trust, which is the one thing you actually need from them.

The honest answer

The fear is real and rarely spoken: is this coming for my job? The platitude ('AI won't replace you') is heard as naive or dishonest and makes it worse. The honest answer names the real change, points to the judgement and relationships that become more valuable, and treats people as adults. That earns the trust the platitude destroys.

3
Building trust instead of mandating use

The instinct, faced with a team that is not adopting, is to mandate: require the tool, track the usage, push harder. This reliably backfires, because the problem was never that people did not know they were allowed to use AI. The problem is that they do not trust it, or fear it, and you cannot mandate your way out of a trust problem. A mandate on top of distrust produces compliance theatre: people use the tool just enough to satisfy the tracking, learn nothing, and trust it even less.

Trust is built the slow way, and there is no shortcut. Let people see it work on low-stakes things first, so they build confidence before the high-stakes ask. Be honest about its limits, the leader who admits where AI fails is far more trusted on where it works than the one who only sells it. Make it safe to be sceptical out loud, because the sceptic who feels heard becomes an ally and the sceptic who feels steamrolled becomes quiet resistance. And answer the job-fear question honestly and repeatedly. None of this is fast, and trying to skip it with a mandate is exactly how rollouts die while the leader wonders why the obviously-good tool is not catching on.

What you walk away with

A straight answer to the job-fear question you can actually say out loud, and a way to build trust that does not rely on hype or mandates. You cannot mandate your way out of a trust problem. Let people see it work small, be honest about limits, make scepticism safe, and answer the fear honestly. Slow, and the only thing that works.

What you are seeing

One of your best CSMs, experienced, trusted by customers, has gone quiet on AI. They use it the bare minimum to satisfy expectations and no more, and lately they seem flatter in general, less engaged. You suspect, correctly, that underneath it they believe AI is making their role pointless, that the craft they built over years is being reduced to checking a machine's output.

The wrong reassurance and what it costs

The tempting move is the warm platitude: "don't be silly, we'll always need great CSMs like you, AI could never do what you do." It feels kind. It is heard as either you not understanding their real fear, or you being dishonest to manage them, and from someone as experienced as they are, it lands as condescending. It costs you the one thing you needed, their belief that you are being straight with them, and it confirms their suspicion that the fear is real and you are papering over it. A strong, disengaging CSM who decides their leader is not honest with them is a renewal risk and a flight risk.

What you actually say and do

You name it honestly: the work is changing, the routine parts genuinely are moving to AI, and you are not going to pretend otherwise. Then you point to the real future and mean it: what makes them exceptional, the judgement, the trust customers place in them, the read of a room, becomes more valuable as the routine work commoditises, and your job is to help them spend more of their time there. Crucially, you change something real, not just say words: you visibly free them from some routine work and put them on the high-judgement accounts or the mentoring role that proves the point. The honest conversation opens the door, but it is the real change that rebuilds the engagement. Words alone, however honest, are still just words to someone who is watching what you do.

Practice check · not scored
Try the judgement call
A strong, experienced CSM has quietly disengaged because they believe AI is reducing their role to checking machine output. You want to re-engage them.

What is the right approach?

An experienced person hears the warm platitude as naive or dishonest, which costs you their trust and confirms their fear. The honest approach names the real change, points to where their value grows, and backs it with a real shift in their work, because words alone are just words to someone watching what you do. A usage target mandates over a trust problem and deepens disengagement, and leaving them alone lets a flight-and-renewal risk quietly grow.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why does a leader's own comfort with AI actively mislead them about their team?

Research shows a wide trust gap, with leaders far more comfortable trusting AI for important decisions than frontline staff. Because you are on the confident side, your own ease leads you to assume the team is there too, so you push a tools solution when the team has a trust problem. It is not that leaders are less skilled, and the gap runs the way described (teams trust it less), which is precisely why a leader's comfort misleads them.

2. Why is the reassurance 'don't worry, AI won't replace you' usually a mistake?

The platitude treats people as children and is heard as either not understanding the real fear or managing them dishonestly, so it costs the trust you need and confirms that the fear is being papered over. The honest answer, naming the real change and pointing to where value grows, is what earns respect. Repetition does not rescue a platitude, and the problem is that it is dishonestly positive, not too negative.

3. Your team is not adopting AI well. Why is mandating use and tracking it likely to backfire?

People are not failing to adopt because they did not know they were allowed, they distrust or fear the tool, and a mandate on top of that produces compliance theatre: minimal use to satisfy tracking, no real learning, even less trust. Trust must be built slowly through low-stakes wins, honesty about limits, and safe scepticism. Mandates do not always work here, tracking is feasible, and training alone does not fix a trust or fear problem.

Score: 0/3 ·

You will leave able to

  • See the gap between how much you trust AI and how much your team does
  • Give an honest answer to the job-fear question without platitudes
  • Build real trust instead of mandating use and hoping

Hands-on exercise

Write the honest answer you would give to a team member who asked, directly, 'is this coming for my job?' Make it true, not comforting. Then pick one nervous or sceptical person on your team and plan a low-stakes way for them to see AI genuinely help, before you ask anything high-stakes of them. Trust is built in that order, small win first, big ask later.

The human element: People can tell the difference between a leader managing their feelings and a leader telling them the truth. The honest, harder conversation is the one that builds the trust you will need for everything else in this course. There is no shortcut, and the attempt to find one is usually a mandate, which is how rollouts quietly die.

If you remember three things

  1. Leaders trust AI far more than their teams do. Your own comfort misleads you, and you end up solving a tools problem when the team has a trust problem.
  2. Answer the job-fear question honestly: the work is changing, judgement and relationships become more valuable, and you will help people move up. The platitude destroys trust.
  3. You cannot mandate your way out of distrust. Build it slowly with small wins, honesty about limits, and safe scepticism.
Part 2 · Build capability · Module 04

Draw the line you can actually defend

A one-page data and tool boundary your team can follow and you can defend upward

16 min 85% to pass

The lesson

1page, plain English

Every CS leader needs a clear answer to one question: what data can go into which tools, and where is the line. Too loose and you risk the incident that sets AI back for everyone. Too tight and your team ignores you and goes back to shadow tools. The job is a line they can follow and you can defend.

1
Why the line is yours to draw, and why both extremes fail

Setting the data line is one of the three things only you can own, and it is the one with the most immediate consequences. The trap is that both ways of getting it wrong feel safe at the time. Draw it too loose, allow real customer data into unapproved tools, and you are one careless paste away from the incident that does not just hurt one account but sets your whole team's AI use back, because after a breach, everything gets locked down. Draw it too tight, ban everything that feels risky, and your team quietly ignores you and goes back to the shadow tools from module 2, except now they hide it harder because the line is unreasonable.

A narrow band, both extremes fail
Too tightteam ignores, shadow use grows
The workable linetight on real risk, loose enough to follow
Too looseone paste from incident
← clamps downopens up →

A line nobody can work within is not a safeguard. It is a fiction that drives risk underground.

So the line has to live in a narrow band: tight enough to prevent the genuine disaster, loose enough that following it is realistic for someone doing the actual job. A line nobody can work within is not a safety measure, it is a fiction that drives risk underground. The skill is finding the band, and writing it so plainly that a CSM under deadline pressure can actually apply it without calling legal.

The point

Both extremes fail and both feel safe. Too loose risks the incident that sets everyone back. Too tight gets ignored and drives shadow use underground. The line has to be tight enough to prevent disaster and loose enough to actually follow. A line nobody can work within is a fiction, not a safeguard.

2
Writing a line people can actually follow

A defensible line is built from a few plain decisions, written so a busy CSM can apply them in seconds, not a policy document only legal can parse. The components below are the whole thing. Keep it to one page in plain language, because a line people cannot remember is a line people cannot follow.

🟢
What is fine, freely

The everyday work that carries no real risk once anonymised: ticket themes, draft structures, general questions. People should feel free here, not nervous, or you have made the whole tool feel forbidden.

🟡
What needs care

Real account material, only in the approved tool, inside the boundary, anonymised where it can be. The 'use the approved tool and anonymise' zone, clearly named so people know the safe path rather than guessing.

🔴
What never goes in

The hard nos: security material, credentials, anything under NDA into an unapproved tool, real customer personal data outside the approved boundary. Short, absolute, and explained, so people follow it because they understand it.

What to do when unsure

The most important line of all: who to ask, and a default of 'if unsure, anonymise or ask, do not guess'. A line without a path for the unsure cases just produces silent wrong guesses.

The format

One page, plain language: what is fine freely, what needs care (approved tool, anonymised), what never goes in, and what to do when unsure. A busy CSM should be able to apply it in seconds without calling legal. If they cannot remember it, they cannot follow it.

3
Defending it without freezing everything

Your line will be tested from two directions, and a good leader holds the middle against both. From above: security and legal may push for a near-total lockdown, because the safest line for them is 'no AI', which is not the safest line for the business that needs the productivity. Your job is to defend a workable line, to show that it genuinely prevents the disaster cases while letting the team do real work, so the answer is not the reflexive ban that just creates shadow use they cannot see.

From the side: your team, especially your best people, will push with 'but this other tool is better'. The wrong responses are the two easy ones, cave (and quietly let the risk in) or clamp down (and drive it underground). The right response is the module 2 move: treat it as a signal, understand why the tool is better, and either get it properly approved or honestly explain the specific risk that keeps it off the line. The phrase that holds the middle is 'show me why it is better and let me see if we can approve it safely', which respects the request, keeps your visibility, and keeps you the one drawing the line rather than the team drawing it for you in the dark.

What you walk away with

A one-page line, plus the ability to defend it both ways: against a security lockdown that would just create invisible shadow use, and against the 'this tool is better' push, by treating it as a signal and either approving safely or explaining the real risk. Cave and you let risk in. Clamp down and you drive it underground. Hold the middle.

The situation

Your strongest CSM is quietly using an unapproved AI tool because it genuinely works better for a key task than your approved one. You have two obvious options in front of you: ban it firmly to enforce the line, or quietly ignore it because they are your best person and clearly know what they are doing. Both feel defensible. Both are wrong.

Why banning it outright is wrong

Banning it outright treats the symptom and ignores the cause. Your best person found a real productivity gain, and a flat ban tells them the line is about control, not safety, which costs you their trust and their honesty. Worse, it does not stop the behaviour, it drives it underground, so now you have the same data risk plus no visibility into it. You have made yourself blind and felt safe doing it, the most dangerous combination there is.

Why ignoring it is also wrong, and what to do instead

Ignoring it is the opposite failure: there may be a genuine data risk (the unapproved tool might be exposing customer data), and 'they are my best person' does not make the data safe, it just means your most capable person is carrying the most risk invisibly. The right move holds the middle: treat the tool choice as the signal it is, understand specifically why it is better, and then either get it approved properly (run it past security, check its data handling) or, if it genuinely cannot be approved, explain the specific risk plainly so they follow the line because they understand it. Either way you keep your visibility, keep their trust, and stay the person drawing the line. The whole skill is refusing both easy answers.

Practice check · not scored
Try the judgement call
Your top CSM is using an unapproved tool because it is better. Your security team's instinct is a blanket ban on all non-approved AI. Your CSM's instinct is that you should just let them use what works.

What is the leadership move that holds the middle?

Both extremes fail: a blanket ban drives real productivity gains and the associated risk underground (you go blind and feel safe), while letting the team use anything lets genuine data risk in unchecked. The middle move treats the tool choice as a signal, then either approves it safely or explains the specific risk so the line is followed out of understanding, keeping both visibility and authority. Avoiding the decision just lets the shadow use and risk continue unmanaged.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why does drawing the data line too tightly fail, even though it feels like the safe choice?

A line nobody can work within is a fiction: people quietly ignore it and go back to shadow tools, except now they hide it harder, so you carry the same risk with less visibility while feeling safe, the worst combination. A tight line is not automatically safest precisely because of this, the failure is about workability not clarity or cost. The line must sit in the band that prevents disaster while staying realistic to follow.

2. What is the single most important component of a usable data line that leaders most often leave out?

Real situations are full of edge cases, and without a path for the unsure cases (who to ask, plus a default of anonymise-or-ask, never guess), people simply guess, usually wrongly and silently. That is the component most often missing. A legal justification makes the line unreadable, an exhaustive tool list dates instantly, and prompt-level tracking is surveillance that does not help someone decide what is safe in the moment.

3. Security wants a blanket ban on all non-approved AI tools. Why should a CS leader push back rather than simply comply?

Security optimises for their risk ('no AI' is safest for them), but that is not safest for a business that needs the productivity, and a blanket ban predictably produces shadow use the leader cannot see. The leader's job is to defend a workable line that genuinely prevents disaster cases while letting the team work. That is not dismissing security's role or claiming to know more than them, and some hard bans are necessary, it is holding the middle between lockdown and a free-for-all.

Score: 0/3 ·

You will leave able to

  • Write a data and tool line your team can follow without a lawyer
  • Defend that line to security and legal without freezing everything
  • Handle the 'this tool is better' push without caving or clamping down

Hands-on exercise

Write your one-page data line now, in plain language: what is fine freely, what needs care (approved tool, anonymised), what never goes in, and exactly what someone does when unsure. Then test it: could a CSM under deadline pressure actually apply it in ten seconds without calling you or legal? If not, it is too complex to be followed, simplify until it can be.

The human element: A good line is an act of respect for your team's intelligence: clear enough to follow, explained well enough to believe in, and reasonable enough that following it is not heroic. People follow lines they understand and ignore lines that treat them as either reckless or stupid.

If you remember three things

  1. Both extremes fail: too loose risks the incident that sets everyone back, too tight drives shadow use underground. The line lives in the workable middle.
  2. One page, plain language: fine freely, needs care, never goes in, and what to do when unsure, the last being the part most often missing.
  3. Defend it both ways: against a security lockdown that creates invisible risk, and against 'this tool is better' by treating it as a signal and approving safely or explaining the real risk.
Part 2 · Build capability · Module 05

Get the whole team using it, not just the keen few

The rollout sequence that moves past the two enthusiasts to real, consistent use

14 min 85% to pass

The lesson

2enthusiasts, then stall

Almost every rollout gets the two natural enthusiasts and then stalls. The hard part is everyone else: the cautious majority and the quietly resistant. This is the practical sequence for getting past the keen few to consistent use, without a clumsy mandate that just creates people pretending to comply.

1
Why rollouts stall after the enthusiasts

Every team has two or three people who would adopt AI no matter what you did, they are curious, they like new tools, they were probably already using it before you said anything. The mistake leaders make is reading these early adopters as proof the rollout is working. They are not. They are the people who never needed convincing, and the adoption numbers they generate hide the fact that everyone else has not moved.

Who you are actually rolling out to
e
e
c
c
c
c
c
c
r
r
2 enthusiasts (adopted anyway)
6 cautious majority (need different things)
2 quiet resisters (trust or fear)

The enthusiasts' numbers hide that no one else moved. Enthusiast tactics on a non-enthusiast audience is how rollouts die.

The real work, and the part that decides whether you get a capable team or a couple of power users in a sea of non-adoption, is everyone after the enthusiasts: the cautious majority who are willing but not driven, and the quietly resistant who have reasons (often the trust and fear reasons from module 3) they are not saying out loud. These people do not respond to the same things the enthusiasts did. Excitement and new features pulled the enthusiasts in. The majority needs something else entirely, and the rollouts that stall are the ones that keep using enthusiast tactics on a non-enthusiast audience.

The point

The two enthusiasts would have adopted anyway, so their usage is not proof the rollout works, it hides that no one else moved. The real work is the cautious majority and the quiet resisters, who need completely different things from what pulled the enthusiasts in. Rollouts stall when leaders keep using enthusiast tactics on everyone else.

2
The sequence that actually spreads

Adoption spreads through people and habit, not through tools and announcements. The sequence below works because it meets the majority where they are, rather than where the enthusiasts were.

🎯
Start with one real, shared workflow

Not 'use AI more', which is too vague to act on. One specific, common task everyone does, done a better way with AI, so success is concrete and copyable. A vague goal produces vague adoption.

🤝
Use peers, not mandates

The cautious majority is moved by a respected colleague saying 'this genuinely saved me time', far more than by a leader's push. Turn an enthusiast into a guide for a few peers. Adoption travels along trust lines, so use them.

👣
Make the first step tiny and safe

The barrier for the cautious is the first use, not the hundredth. A small, low-stakes, supported first win builds the confidence that makes the next step easy. Big asks freeze cautious people.

👈
Address the resisters' real reason

Quiet resistance is usually trust or fear (module 3), not laziness. Pushing harder deepens it. Find the real reason and address that, and the resister often becomes the most thorough adopter, because they needed to believe it first.

The principle

Adoption spreads through people and habit, not tools and announcements. Start with one concrete shared workflow, spread it through respected peers rather than mandates, make the first step tiny and safe, and address the resisters' real reason instead of pushing. Meet the majority where they are, not where the enthusiasts were.

3
Why the mandate is the great destroyer

When adoption stalls, the single most tempting move is to mandate: require the tool, set a usage target, track it, and push. It is tempting because it feels like decisive leadership and it produces a number that goes up. It is also the move most likely to kill real adoption, and understanding why is the heart of this module.

A mandate on a willing person is unnecessary, they were going to adopt anyway. A mandate on a cautious or resistant person produces compliance theatre: they use the tool exactly enough to hit the target and satisfy the tracking, learn nothing real, and quietly resent it, which hardens the very resistance you were trying to overcome. So the usage number goes up while genuine capability flatlines or falls, and you have manufactured a metric that lies to you. The honest read of a mandate is that it converts a trust-and-habit problem, which is solvable slowly, into a compliance problem, which is not solvable at all because it is the wrong frame. The leaders who build real capability resist the mandate precisely when it is most tempting, which is when things have stalled.

What you walk away with

A rollout plan that does not die after the enthusiasts: one concrete workflow, peer-led spread, tiny safe first steps, and the resisters' real reasons addressed. And the discipline to resist the mandate when it is most tempting, because it manufactures a usage number that lies while hardening the resistance underneath.

What the dashboard shows

Three months into your rollout, the usage numbers look healthy, plenty of AI activity logged across the team, and you are inclined to call it a success. Then you look closer and realise almost all of that activity comes from the same two people who were enthusiastically using AI before you launched anything. The other eight on your team have barely moved. The number is real, but it is the enthusiasts' number, and it is hiding a stalled rollout.

What a mandate would do here

The tempting fix, faced with eight people who have not moved, is a usage mandate: everyone must log a minimum amount of AI use per week, tracked. Watch what it produces. The eight cautious and resistant people start doing the minimum to hit the target, copy-pasting a token task or two, learning nothing, and quietly resenting being measured on it. Your dashboard number jumps, which feels like success. But genuine capability has not moved, and the resentment has hardened the resistance, so you are now further from real adoption than before, while your metric tells you the opposite. You have manufactured a lie and made the underlying problem worse.

The genuine unlock

The real unlock is not a mandate, it is meeting the eight where they are. Pick one concrete workflow they all do, have one of the two enthusiasts (a respected peer, not you) show two or three of them how it genuinely saved time, make their first attempt tiny and supported, and for anyone still resisting, find out whether it is trust or fear (module 3) and address that rather than pushing. This is slower than a mandate and produces a worse-looking number next month, and it is the only thing that produces real capability. The whole lesson: the honest slow path beats the dishonest fast number, and the mandate is the dishonest fast number.

Practice check · not scored
Try the judgement call
Your AI usage numbers look healthy, but you discover the activity is almost entirely from the two people who were already keen before the rollout. The other eight have barely engaged.

What is the genuine unlock for the eight?

The eight are the cautious majority and quiet resisters, who need peer-led, concrete, low-stakes on-ramps and, where it is trust or fear, that real reason addressed, not enthusiast tactics. A mandate produces compliance theatre: the number rises while capability does not and resentment hardens. Accepting the enthusiasts' number as success leads the fiction, and replacing the eight treats a normal adoption curve as a hiring problem.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why are healthy-looking adoption numbers, driven mostly by the early enthusiasts, a trap for a leader?

Early enthusiasts never needed convincing, so their (real) activity can make a stalled rollout look successful by masking that everyone else stayed put. The number is genuine but unrepresentative. It is not that the usage is fake or that high numbers are inherently bad, and the issue is certainly not that enthusiasts use AI too much, it is that their adoption is not evidence the rest of the team has moved.

2. What moves the cautious majority to adopt, given that excitement and new features are what pulled the enthusiasts in?

The majority needs something different from the enthusiasts: concreteness (one real workflow), social proof from a trusted peer, and a small, safe, supported first step, because their barrier is the first use, not the hundredth. Enthusiast tactics fail on them, a directive triggers compliance theatre, and withholding help freezes cautious people rather than activating them.

3. Why is mandating AI usage the move most likely to kill real adoption, precisely when a rollout has stalled?

A mandate converts a solvable trust-and-habit problem into an unsolvable compliance problem: cautious and resistant people do the minimum to satisfy tracking, learn nothing, and resent it, hardening resistance while the number rises and lies. The danger is not enforcement difficulty or measurement accuracy, and it certainly does not work too well, it manufactures a misleading metric while making genuine capability worse.

Score: 0/3 ·

You will leave able to

  • Sequence a rollout that keeps going past the keen few
  • Bring along the cautious and the quietly resistant without forcing it
  • Avoid the mandate that creates compliance theatre instead of real use

Hands-on exercise

Map your team onto the curve: who are your two enthusiasts, who is the cautious majority, who is quietly resisting. Then design the next step for the majority, not the enthusiasts: one concrete shared workflow, which respected peer will show it, and what the tiny first step is. For each resister, write down what you think their real reason is, and resist the urge to put a number on anyone.

The human element: Real adoption is a change in habit, and habits change through trust and small wins, the same way they always have. The mandate is the shortcut that is not a shortcut, it buys a number and costs you the thing the number was supposed to represent.

If you remember three things

  1. The two enthusiasts would have adopted anyway. Their usage hides whether the rest of the team has actually moved.
  2. Spread adoption through people and habit: one concrete workflow, peer-led, tiny safe first steps, and the resisters' real reasons addressed.
  3. Resist the mandate when it is most tempting. It manufactures a usage number that lies while hardening the resistance underneath.
Part 2 · Build capability · Module 06

Get consistent quality without crushing your people

Raising the floor without lowering the ceiling, and coaching the over-reliant

16 min 85% to pass

The lesson

raise floorkeep ceiling

Once the team is using AI, a new problem appears: quality swings wildly from person to person, and the work starts to sound the same. This module is raising the floor without lowering the ceiling, a shared standard so customers get consistent quality, while good people keep their judgement and voice.

1
The two new problems success creates

Adoption succeeding does not end your quality job, it changes it. Two new problems appear, and they pull in opposite directions, which is what makes this hard. First, quality variance: some people use AI to produce excellent work, others produce confident rubbish, and the customer experience now swings depending on who they happen to deal with. Second, sameness: as everyone leans on the same tools, the work starts to converge on a generic AI voice, and your team's output begins to sound like every other vendor's.

The instinct that solves the first problem makes the second worse. Faced with variance, leaders reach for rigid templates and mandated prompts to force consistency, which does raise the floor, but it also flattens the ceiling, crushing the judgement and voice of your best people and accelerating the sameness. So the real task is the harder both-at-once: lift the weakest work to a reliable standard without dragging the strongest work down to a template. Raise the floor, keep the ceiling.

The point

Success creates two new problems that pull opposite ways: quality variance (some produce excellent work, some confident rubbish) and sameness (everyone converging on a generic AI voice). The rigid-template fix raises the floor but flattens the ceiling. The real job is lifting the weakest work without crushing the strongest. Raise the floor, keep the ceiling.

2
A shared standard that is not a straitjacket

The way to get consistency without uniformity is to standardise the outcome, not the method. Set a clear bar for what good output must achieve, every customer email is verified for accuracy, sounds like a human wrote it, and contains a specific next step, without dictating the exact words or prompt that gets there. That lifts the floor (the weak work now has a bar to clear) while leaving the ceiling open (strong people hit the bar their own way and keep their voice).

Reviewing AI-assisted work fairly follows the same principle: you judge the output against the standard, not whether they used AI or which tool or prompt. The question is never 'did you use AI for this?' but 'does this meet the bar, is it accurate, does it sound like us, does it serve the customer?'. And you spread quality the productive way, by taking what your best people genuinely do and making it available to everyone as a guide, not a mandate, so the strong lift the weak rather than being levelled down to them. The standard is the floor everyone clears. How they clear it is theirs.

The principle

Standardise the outcome, not the method. Set a clear bar for what good output must achieve and judge work against that, not against whether or which AI was used. Spread what your best people do as a guide, not a mandate. The standard is the floor everyone clears. How they clear it stays theirs, which keeps the ceiling and the voice.

3
Coaching the over-reliant, before it costs a renewal

The hardest individual problem AI creates for a leader is the person whose numbers look fine while their underlying judgement quietly erodes. They lean on AI for everything, their output passes review, and on the dashboard they look like a strong adopter. But they have stopped being able to do the thinking themselves, and the day a call goes off-script, or a situation arrives that the AI prep did not cover, they freeze, because the skill that used to handle it has faded from disuse. This is the 'quietly at risk' pattern from module 2, and it is dangerous precisely because it is invisible until the moment it costs you.

Coaching it is delicate, because the person is, by the numbers, succeeding, so a clumsy intervention feels like punishment for doing well and demotivates a good performer. The move is to frame it as protecting their value, not correcting a failure: name honestly that the routine work is well-handled and that the judgement, the live, unscripted, human part, is what makes them valuable and is the thing to keep sharp. Then build deliberate practice of that judgement back in: the occasional account worked without AI prep, the live scenario rehearsed, the decision made and explained before the tool is consulted. You are not asking them to use AI less in general. You are making sure the muscle underneath does not waste away, framed as investment in them, because it is.

What you walk away with

A shared quality bar that lifts the floor without flattening the ceiling, fair output-based review, and a way to coach the over-reliant CSM, the one whose numbers look fine while their judgement erodes invisibly. Coach it as protecting their value, not correcting a failure, and rebuild the judgement muscle with deliberate practice before the off-script moment exposes it.

The situation

One of your CSMs has great numbers and clean reviews, a model adopter on paper. Then you sit in on a call that goes somewhere their AI prep did not anticipate, and they freeze, visibly lost without the script. Afterwards you realise the pattern: they cannot really handle an escalation or an unscripted moment any more without running AI prep first, and when reality outruns the prep, the judgement that used to carry them is not there. It has faded from disuse, hidden the whole time behind good metrics.

The clumsy intervention and what it costs

The tempting move is to address it as a performance problem: tell them they are too dependent on AI, that they need to cut back, perhaps flag it in a review. To someone whose numbers are genuinely good, this lands as being punished for succeeding, it is confusing and demotivating, and it frames the very tool you spent three modules getting them to adopt as the thing they did wrong. You risk souring a strong performer on the whole capability and on you, and you still have not actually rebuilt the missing judgement, you have just made them anxious about a tool.

Coaching it the right way

Frame it as protecting what makes them valuable, because that is true. Acknowledge honestly that they handle the routine work really well, and that the thing that makes them exceptional, the live judgement, the read of a hard moment, is exactly what froze on that call and is worth keeping sharp. Then build deliberate practice of it back in, framed as investment, not correction: rehearse the off-script scenarios, occasionally work an escalation through their own judgement first before consulting the tool, make the decision and explain the reasoning unaided. You are not telling them to use AI less across the board, you are making sure the muscle underneath the prep stays strong, so the next off-script moment finds them ready. Done this way, a good performer feels invested in, not punished, and the renewal-risk that was hiding behind their metrics quietly closes.

Practice check · not scored
Try the judgement call
A CSM with strong metrics and clean reviews freezes when a call goes off-script, because they have become unable to handle escalations without AI prep first. You need to coach this without demotivating them.

What is the right approach?

The person is succeeding by the numbers, so a correction frames AI (which you worked to get them to adopt) as a failure and demotivates a strong performer without rebuilding the missing judgement. The right move frames it as investment in their value and rebuilds the judgement muscle with deliberate practice. Doing nothing leaves a renewal risk hidden behind good metrics, and cutting their access is a punitive blunt instrument that breeds resentment rather than rebuilding skill.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why does the rigid-template fix for quality variance make the second problem, sameness, worse?

Mandated templates and prompts do raise the weakest work to a standard, but by dictating the method they also strip the strongest people of their judgement and voice, pushing everyone toward the same generic output, which is exactly the sameness problem. They do affect sameness (worsening it), they do help the weakest work, and difficulty is not the issue, the issue is they flatten the ceiling while raising the floor.

2. What is the key principle for getting consistent quality without uniformity?

Standardising the outcome sets a floor everyone must clear (accurate, human-sounding, useful) while leaving them free to clear it their own way, which preserves the ceiling and individual voice. Standardising prompts is the rigid-template trap that flattens the ceiling, judging by whether or which AI was used is the wrong question (you judge the output against the bar), and no shared standard at all leaves the variance problem unsolved.

3. Why is the over-reliant CSM, whose metrics look fine, such a dangerous case for a leader?

The danger is precisely the invisibility: good metrics and clean reviews hide that the underlying judgement has faded, so the problem only surfaces in the unscripted moment the prep did not cover, potentially in front of a customer at renewal time. They are not obviously underperforming, they use AI heavily not too little, the passing reviews are what mask the risk rather than removing it, and it does not resolve itself, it requires deliberate coaching.

Score: 0/3 ·

You will leave able to

  • Set a shared quality standard without forcing everyone into the same template
  • Review AI-assisted work fairly, against the outcome rather than the tool used
  • Coach the over-reliant CSM before their fading judgement costs a renewal

Hands-on exercise

Write your team's quality bar as outcomes, not methods: what must every piece of customer-facing work achieve, regardless of how it was produced. Then identify anyone who might be 'quietly at risk', strong numbers, possible eroding judgement, and plan one piece of deliberate unscripted practice for them, framed as protecting their value, not correcting a fault.

The human element: The goal is a team that is more capable, not a team that is more uniform. The leader who protects each person's judgement while raising the shared standard builds something resilient. The one who templates everything builds a fragile team that produces identical work and cannot cope the day the script runs out.

If you remember three things

  1. Success creates two opposing problems: quality variance and sameness. The rigid-template fix solves the first by worsening the second.
  2. Standardise the outcome, not the method. Set a clear bar, judge work against it not against whether AI was used, and spread your best people's approach as a guide.
  3. Coach the over-reliant CSM as protecting their value, not correcting a failure. Rebuild eroding judgement with deliberate practice before an off-script moment exposes it.
Part 3 · Prove and protect · Module 07

Prove the value in language your leadership respects

Measuring what matters and reporting it honestly, so you keep the trust and the investment

14 min 85% to pass

The lesson

1slide that matters

Sooner or later your VP asks the real question: what has all this AI actually done for us? A lot of leaders cannot answer well, they have activity numbers but no proof of value. This module is measuring what actually matters and reporting it honestly, so you keep the trust and the investment.

1
Activity is not value, and leadership knows it

The trap that catches most CS leaders here is reporting activity and calling it value. 'The team ran 4,000 AI queries this quarter' and 'adoption is at 90 per cent' are activity numbers, they measure that the tool is being used, not that anything good came of it. A sharp VP hears these and asks the obvious follow-up, 'and what did that get us?', and if you do not have an answer, you have just revealed that you measured the input and never checked the output.

Value is the change in business outcomes, not the volume of tool use. Did the team genuinely get time back, and was it reinvested in something that mattered? Were risks caught earlier than they would have been? Were renewals protected or expansions opened that can be traced, even loosely, to the capability you built? These are harder to measure than query counts, which is exactly why most leaders fall back on activity, and exactly why the leader who can speak to real outcomes stands out. Leadership funds value, not usage, so usage numbers do not protect your investment, they just prove you spent the budget.

The point

Activity (queries run, adoption percentage) measures that the tool is used, not that anything good came of it, and a sharp VP will ask 'what did that get us?'. Value is the change in outcomes: time genuinely reinvested, risks caught earlier, renewals protected. Leadership funds value, not usage. Reporting activity proves you spent the budget, not that it worked.

2
The metrics that actually mean something

A small set of honest outcome measures beats a dashboard of activity. The point is not precision, you will rarely get clean attribution, it is connecting the capability to things leadership already cares about, in their language.

⏱️
Time reinvested, not just saved

Hours saved is half the story and the weak half. The strong half is where those hours went: more time on at-risk accounts, more face time with key customers. 'Saved' is a cost story, 'reinvested into X' is a value story.

🚨
Risk caught earlier

Cases where the team spotted a churn signal or an issue sooner than they would have, because the capability surfaced it. Earlier detection is worth real money and leadership understands it instantly.

💰
Outcomes you can trace

Renewals protected, expansions opened, that you can connect, even loosely and honestly, to the capability. Loose-but-honest attribution beats precise-but-meaningless activity every time.

📉
What is not working

Counterintuitively a value metric: naming where AI has not helped, or where adoption is genuinely stuck, makes every other number you report believable. The honest negative is what earns trust in the positives.

The framing

A few honest outcome measures beat a dashboard of activity: time reinvested (not just saved), risk caught earlier, traceable renewals or expansions, and an honest account of what is not working. Loose-but-honest attribution beats precise-but-meaningless counts. Speak in the outcomes leadership already cares about, in their language.

3
Reporting honestly is the long game

The pressure, when asked to prove value, is to oversell, to claim more than you can support, attribute every good thing to AI, and hide what is not working. It buys a good meeting and costs you everything after it, because the first time a claim does not hold up, every future number you report is discounted. Overclaiming is a loan against your credibility at a punishing interest rate.

The honest report is the one that compounds. Include what is not working, and name the limits of your attribution ('we cannot prove this renewal was the AI, but the early signal that flagged it came from the new process'). This feels weaker in the moment and is far stronger over time, because a leader who reports the negative honestly is believed on the positive, and a leader who only ever reports wins is quietly assumed to be spinning. The goal is not to win one meeting with an inflated number. It is to become the person whose numbers leadership trusts, which is what actually protects the investment across many meetings, and is worth far more than any single impressive but fragile claim.

What you walk away with

A short, honest set of outcome metrics and a one-slide answer to 'what has AI done for us'. Resist overclaiming: it buys one meeting and discredits every future number. Report what is not working and name your attribution limits. The honest negative is what makes you the leader whose numbers are trusted, which is what actually protects the investment.

The weak slide

Your VP asks what AI has delivered this quarter. The tempting slide is a wall of activity: 90 per cent adoption, 4,000 queries, every CSM 'using AI daily', and a confident headline like 'AI saved the team 600 hours and drove this quarter's renewals'. It looks impressive for about ten seconds, until the VP asks 'how do you know it drove the renewals?' and 'what did the 600 saved hours actually produce?', and you do not have an answer, at which point the whole slide deflates and your credibility with it.

What goes on the strong slide, and what comes off

The strong slide drops the activity numbers and the unsupported headline. On it: the team recovered roughly six hours each per week (stated as an estimate, not a precise claim), and here is specifically where those hours went, more time on the at-risk accounts, which connects to two named saves this quarter where the early warning came from the new process. It also, deliberately, includes one honest line about what is not working: adoption is genuinely strong on call prep but has stalled on risk analysis, and here is the plan. The tempting number to leave off is the big 'AI drove the renewals' attribution, because you cannot support it and claiming it puts every other number at risk.

Why the honest slide wins the longer game

The weak slide might win a single meeting if no one probes, but it is fragile, one good question collapses it, and even if it survives, the inflated claim becomes a debt that comes due later. The strong slide is more modest in the moment and far more durable: the VP can see exactly what is real, the honest 'not working' line makes the positive claims believable rather than suspicious, and you have established yourself as the leader whose numbers can be trusted. That trust is what actually protects the investment over the next year of meetings, which is the real game, not the single slide. The honest, slightly less impressive answer is the one that keeps the funding.

Practice check · not scored
Try the judgement call
Your VP asks what AI has actually delivered and you get one slide. You have strong activity numbers and a genuine but hard-to-attribute sense that it has helped.

What belongs on the slide, and what should you leave off?

Leadership funds value, not activity, so the slide should carry honest outcomes (hours reinvested into named work, earlier risk catches, traceable results) plus one honest 'not working' line that makes the rest believable, while leaving off the unprovable big attribution that would collapse under a single question and discredit everything else. Adoption and query counts are activity, the unprovable renewal claim is the exact thing to omit, and reporting only positives reads as spin.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why is reporting '90% adoption and 4,000 queries' a weak answer to 'what has AI done for us'?

Activity measures usage, not value, so it invites the obvious follow-up, 'and what did that get us?', which exposes that you measured the input and never checked the output. The problem is not accuracy or that high adoption is bad, and leadership does care about adoption as a means, but what they fund is the outcome it produces, so value, not activity, is what answers the real question.

2. What makes 'time reinvested' a stronger value metric than 'time saved'?

Hours saved is only half the story and the weaker half, a cost reduction, whereas showing where those hours went (more time on at-risk accounts, more face time) turns it into a value story tied to outcomes leadership funds. The two are not interchangeable, 'saved' is not more impressive, reinvested time is generally harder not easier to quantify, and saved time can be estimated, the point is that reinvestment is the part that demonstrates value.

3. Why does including what is NOT working actually strengthen your report to leadership?

Reporting the honest negative is itself a value move: a leader who names what is not working is believed on what is, while one who reports only wins is assumed to be spinning, which discredits even the real successes. It is not about giving leadership reasons to cut budget or cynically lowering expectations, it is that honesty is what makes you the leader whose numbers are trusted, which protects the investment over time.

Score: 0/3 ·

You will leave able to

  • Choose metrics that show real value, not vanity activity numbers
  • Report honestly, including the parts that are not working
  • Make the case for continued investment without overclaiming

Hands-on exercise

Build your one slide now. Drop every pure activity number. Write down: where did recovered time actually go, what risk was caught earlier because of the capability, what outcome can you trace (even loosely and honestly) to it, and one true thing that is not working. Then find the tempting overclaim you were going to make, and cut it.

The human element: Trust with your leadership is built the same way trust with your team is: by being the person who tells them the truth, including the inconvenient parts. The honest reporter keeps the investment across many meetings. The overclaimer wins one and is discounted ever after.

If you remember three things

  1. Activity (queries, adoption %) is not value. Leadership funds outcomes, and a sharp VP will ask what the activity actually got them.
  2. Report a few honest outcomes: time reinvested into specific work, risk caught earlier, traceable results, and what is not working.
  3. Resist overclaiming. It wins one meeting and discredits every future number. The honest report is what makes you the leader whose numbers are trusted.
Part 3 · Prove and protect · Module 08

Be ready for the governance question before it is asked

Accountability and oversight so an AI-assisted decision can always be explained

14 min 85% to pass

The lesson

78% are not audit-ready

A lot of leaders quietly admit they could not pass an audit of their AI use if asked today. You do not want to be one of them. This module is getting ahead of it: knowing who is accountable for what, and being able to show your team's use is deliberate and controlled, not accidental.

1
The question is coming, and 'we're careful' is not an answer

At some point, someone with authority, security, legal, a regulator, your own leadership after an incident somewhere else, is going to ask how your team's AI use is governed. Most CS leaders are not ready for this. They have a vague sense that the team is sensible, but no actual answer to 'who is accountable for AI-assisted decisions, and how do you control this?'. 'We're careful' is not an answer, it is the absence of one, and it is exactly what fails the moment the question is asked seriously.

Being ready is not about heavy bureaucracy, and the fear that it must be is why leaders avoid it until forced. It is about being able to show that your team's AI use is deliberate and controlled rather than accidental and hopeful. The difference between a quiet risk sitting on your team and a capability you can stand behind is whether you can answer the governance question with evidence instead of reassurance. And the only good time to build that answer is before you are asked, because building it under the pressure of an actual incident or audit is far worse.

The point

Someone with authority will eventually ask how your team's AI use is governed, and 'we're careful' is the absence of an answer. Being ready is not heavy bureaucracy, it is being able to show your team's use is deliberate and controlled, not accidental. Build that answer before you are asked, because building it during an incident is far worse.

2
The three things that make you ready

Audit-readiness for a CS team is lighter than the word suggests. Three things cover most of it, and none requires a compliance department, just deliberate clarity you can produce on request.

👤
Clear accountability

Who is responsible for AI-assisted decisions, the CSM who made the call, owning the output regardless of the tool. 'The AI decided' is never an acceptable answer, a named human always owns it. Establish that as a principle the team understands.

📋
A written, followed line

Your module-4 data line, written down, and genuinely followed. A line you can produce on request, that the team actually applies, is most of what 'how do you control this' asks for. An unwritten or ignored line is no control at all.

🔍
Light oversight

Some regular, proportionate way you actually check, spot-reviewing AI-assisted work, knowing where the risk concentrates, catching a problem yourself before it becomes an incident. Oversight you do, not oversight you claim.

The standard

Three things, none requiring a compliance department: clear accountability (a named human always owns the decision, never 'the AI did it'), a written line that is genuinely followed, and light, real oversight you actually perform. Together these let you answer 'how do you control this' with evidence rather than hope.

3
Catching the problem before it becomes the incident

The deepest purpose of oversight is not to satisfy an auditor, it is to be the person who catches the problem on your own team before it becomes the incident that triggers the audit in the first place. A leader with real, light-touch oversight notices the CSM relying on AI past their judgement, the account where something does not look right, the drift toward an unapproved tool, while it is still small and fixable. A leader with no oversight finds out the same way everyone else does: when it has already gone wrong and become someone else's question.

This is the difference between AI being a quiet, unexamined risk on your team and being a capability you can genuinely stand behind. The standing-behind is not a posture, it is earned by actually looking: knowing where the risk concentrates, checking proportionately, and acting on what you find early. When the governance question comes, the leader who has been doing this can answer it calmly with evidence, and more importantly has probably already prevented the incident that would have made the question hostile. The readiness and the prevention are the same practice, looking at your own team's AI use deliberately, which almost no leader does until forced, and which is the entire point of this module.

What you walk away with

A simple accountability and oversight setup so you are never caught out by the governance question. Its deepest value is not passing an audit, it is catching the problem early, before it becomes the incident that triggers one. Readiness and prevention are the same practice: deliberately looking at your own team's AI use, which almost no leader does until forced.

The challenge

Months after the fact, a decision one of your CSMs made, partly informed by AI analysis, is questioned, perhaps a renewal call that went wrong, perhaps a risk that was missed. Someone with authority asks you to explain how that decision was made, who was responsible for it, and how your team's AI use is controlled in general. This is the governance question, arriving in its most pointed form: about a specific decision that is now under scrutiny.

Whether your current setup could answer it

Ask yourself honestly whether you could answer right now. Could you say who owned that decision, a named human, not 'the AI suggested it'? Could you produce the line that governed what data went into the tool, and show it was followed? Could you show any oversight that should have caught a problem, or explain why this one slipped through? For most leaders the honest answer today is no, they would be reconstructing it in a panic, and the gaps would show. That panic, after the fact, under scrutiny, is the worst possible time to discover you were never ready.

What would have to be true for you to answer it well

Three things, all buildable in advance and none heavy. First, accountability was clear all along: that CSM owned the decision, used the AI as input, and that was understood by everyone, so 'who was responsible' has a clean answer. Second, the written data line existed and was followed, so 'how is this controlled' is answered by producing it and showing the practice. Third, you had light oversight, so you can either show it should have caught this and will be tightened, or honestly show this was a genuine edge case outside reasonable controls, which is a defensible answer when the controls are real. The leader who built these three calmly answers a hostile question with evidence. The leader who did not is exposed, and the lesson is that the only time to build the answer was before the question, which is now.

Practice check · not scored
Try the judgement call
A regulator-style review asks you to explain how a specific AI-assisted decision on your team was made, who owned it, and how your team's AI use is controlled. You want to know whether you are ready.

What three things determine whether you can answer well?

Audit-readiness for a CS team rests on three light things: clear human accountability ('the AI did it' is never acceptable), a written line that is genuinely followed, and proportionate oversight you actually do, which together answer the question with evidence. It does not require a compliance department or bureaucracy, no tool is certified never to err (the human owns the decision precisely because of that), and 'we're careful' is the absence of an answer, not an answer.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why is 'we're careful' a failing answer to 'how is your team's AI use governed'?

Governance questions want evidence of deliberate control, and 'we're careful' provides none, it is a vague sense in place of an actual answer, which is exactly what fails under serious scrutiny. Care matters but is not the same as demonstrable control, the answer is genuinely insufficient (not sufficient) for a real review, and the problem is that it offers no evidence, not that it admits carelessness.

2. What is the deepest value of having real oversight of your team's AI use?

Oversight's deepest purpose is prevention: a leader actually looking catches the eroding judgement, the off account, the drift to an unapproved tool while it is still small, often preventing the incident entirely. Readiness and prevention are the same practice. It is not about paperwork volume or shifting blame (accountability is about a human owning decisions, not scapegoating), and its value is precisely before an incident, not only after.

3. When an AI-assisted decision is challenged, why must a named human always own it rather than 'the AI'?

Clear accountability means a named human owns every AI-assisted decision, using the tool as input rather than as the decider, because 'the AI decided' is never an acceptable governance answer and there must always be a responsible person. The vendor is not the accountable party for your team's decisions, blaming the AI is the opposite of safe (it signals no control), and AI is not infallible, which is exactly why human ownership is required.

Score: 0/3 ·

You will leave able to

  • Show who is accountable for AI-assisted decisions on your team
  • Answer the 'how do you control this' question with evidence, not hope
  • Catch a problem through your own oversight before it becomes an incident

Hands-on exercise

Run the test on yourself today: if asked right now to explain how a specific AI-assisted decision on your team was made and controlled, could you? Write down your honest answer for the three pillars: is accountability clear (a named human owns decisions), is your data line written and followed, and do you have any real oversight. Each 'no' is a gap to close before the question arrives.

The human element: Governance done well is not bureaucracy, it is care made visible and accountable. The leader who looks at their own team's AI use deliberately protects both their customers and their people, and is never the one caught reconstructing an answer in a panic after something has already gone wrong.

If you remember three things

  1. The governance question is coming, and 'we're careful' is the absence of an answer. Build a real one before you are asked, not during an incident.
  2. Three light things make you ready: clear human accountability, a written line that is followed, and real oversight you actually perform.
  3. The deepest value of oversight is prevention, catching the problem early. Readiness and preventing the incident are the same practice.
Part 3 · Prove and protect · Module 09

Lead the change without losing the human core of the job

Protecting what must stay human, and leading people through real anxiety honestly

16 min 85% to pass

The lesson

the soulof the job

The last piece matters most over time. As more work gets automated, the leader's job is to protect the parts of Customer Success that must stay human, the judgement, the relationships, the craft, and to lead a team through the genuine anxiety of all this without pretending it away.

1
What must stay human, on purpose

Everything in this course has been about building AI capability. This final module is about the thing that capability must not be allowed to erode, because if it does, you have built an efficient team that has lost the reason customers stayed with it. Some parts of Customer Success are not inefficiencies to be optimised away, they are the actual product: the judgement that reads a situation no dashboard captures, the relationship a customer trusts when things go wrong, the craft of a hard conversation handled well. These must stay human not out of sentiment but because they are what your team is genuinely for.

The leader's job is to protect these deliberately, because nothing else will. Left alone, the logic of efficiency quietly eats them: every human touch looks like a cost to optimise, until you have automated the relationship out of a relationship business. Defining what stays human, and defending it against your own efficiency drive, is a leadership act. It is also, increasingly, a competitive one: in a market where everyone has the same AI, the team that kept its human core is the one customers can tell apart from the rest.

The point

Some parts of CS are not inefficiencies to optimise away, they are the actual product: the judgement, the relationships, the craft. Left to the logic of efficiency, they get quietly eaten. The leader's job is to protect them deliberately, which is both a leadership act and, in a market where everyone has the same AI, a competitive one.

2
Leading through the anxiety honestly

Your team is carrying real anxiety about all this, the job fear from module 3, now deepened by months of watching the work change. The temptation, again, is the comforting lie: nothing is really changing, you are all safe, do not worry. It fails for the same reason it failed in module 3, people can tell, and it costs you the trust you need to lead them through what is actually a hard transition.

Leading through it honestly means three things held together. Name the change truthfully: yes, the work is shifting, the routine parts are increasingly automated, and that is real. Point to the durable core: what makes you valuable, the human judgement and relationships, is becoming more important, not less, and that is also real. And change something to prove it, not just say it: visibly invest your team's freed-up time in the human, high-judgement work, so they experience the future you are describing rather than just hearing about it. The anxiety does not respond to reassurance, it responds to a leader who is honest about the change and demonstrably moving the team toward the valuable part of it. That is what people will follow.

The honest path

The team's anxiety is real and the comforting lie fails, because people can tell. Lead through it by holding three things together: name the change truthfully, point to the durable human core that grows more valuable, and prove it by visibly moving the team's freed time toward that work. Anxiety responds to honest movement, not reassurance.

3
What good looks like, and being the leader people follow

Pulling the whole course together, here is what good looks like for a CS team in an AI-heavy world: AI does the assembly and the routine at a consistently high standard, the team's human judgement sits at the centre of every decision that matters, the relationships are deeper because people have more time for them rather than less, and everyone, including the nervous and the once-resistant, understands where they are genuinely valuable and is being moved toward it. That is not a team that resisted AI, and not a team that surrendered to it. It is a team that used it to become more of what it was for.

Being the leader who gets a team there is the sum of everything in this course: you saw the real picture, closed the trust gap honestly, drew a defensible line, brought everyone along without mandates, held quality without crushing people, proved value truthfully, stayed ready for the hard questions, and protected the human core through the change. None of it was about being the best with the tool. All of it was about leading people through a genuine transition with honesty and judgement, which is the oldest leadership job there is, wearing new clothes. The leaders people follow through this are not the most technically dazzling. They are the ones who were honest about the change and clearly had their team's interests at the centre of it. Be that one.

What you walk away with

A clear definition of what good looks like, a team that used AI to become more of what it was for, judgement at the centre, relationships deeper, everyone moved toward where they are valuable, and a way to lead people there honestly. The whole course in one line: leading people through a real transition with honesty and judgement, which is the oldest leadership job there is.

The fear, stated plainly

A capable team member comes to you, or you sense it without them saying it, with a specific version of the job fear: not that they will be fired, but that the job itself is being hollowed out, that they are being turned into someone who just checks the AI's output, a quality-control step on a machine, rather than the skilled CSM they trained to be. It is a real fear, and unlike a vague worry it is precise, which means a vague reassurance will obviously miss it.

What not to say, and why

The failing responses are the two easy ones. The comforting lie, 'don't worry, you're so much more than that, nothing's really changing', misses the precise fear and is heard as not listening, because they have correctly noticed that something is changing. The dismissive version, 'this is just the way things are going, adapt', is honest about the change but abandons them in it, confirming exactly the fear that they are now disposable plumbing. Both lose the person: one by not being honest, the other by not caring. The fear is precise and deserves a response that is both honest and invested.

What you say, and what you actually change

You agree with the true part of their fear, which disarms it: yes, if the job became only checking AI output, that would be a worse job, and you are not going to pretend that risk is not real. Then you reframe around the durable core honestly: the checking is not the job, it is the floor, and the part that makes them valuable, the judgement about which of the AI's reads to trust, the relationship that turns a flagged risk into a saved account, the hard conversation no model can have, is becoming the larger part of the job, not the smaller one. And then, crucially, you change something real to prove it: you visibly move them toward more of that human, high-judgement work and less pure output-processing, so they live the reframe instead of just hearing it. The honesty opens the door, the real change walks them through it, and a person who was quietly deciding the job had lost its meaning re-engages because their leader was straight with them and then did something about it. That, in one conversation, is the whole module.

Practice check · not scored
Try the judgement call
A capable CSM fears the job is being hollowed out, that they are becoming a quality-control step that just checks AI output rather than a skilled professional. You want to lead them through this honestly.

What is the right response?

The fear is precise, so it needs a response that is both honest and invested: agree the true part to disarm it, reframe around the durable human core that is growing, and prove it with a real change in their work so they live the future rather than hear about it. The comforting lie misses the fear and is heard as not listening, 'just adapt' abandons them and confirms the fear, and cutting AI use retreats from the change rather than leading through it.

You've reached the knowledge check. Real leadership situations. Work through all of them before answers are revealed. You need 85% to pass.

Knowledge check

3 questions · 85% to pass

1. Why must a leader protect the human parts of CS deliberately, rather than assuming they will survive on their own?

The human core, judgement, relationships, craft, is the actual product, but the unchecked logic of efficiency treats every human touch as a cost and erodes it until a relationship business has automated away its relationships. So it must be protected deliberately. It is not driven by an explicit customer demand, the human parts are the product rather than inefficiencies, and they will not survive on their own, which is the whole reason leadership protection is required.

2. Why does the reassurance 'nothing is really changing, you're all safe' fail when leading a team through AI anxiety?

The comforting lie fails because the change is real and visible, so denying it is heard as either not understanding or not being straight, which costs trust, exactly as in module 3. Confidence does not rescue an untrue claim, the team does not want doom either, and reassurance is not always wrong, the problem is specifically reassurance that contradicts what people can plainly see. Honest acknowledgement plus a real reframe is what works.

3. What does 'good' look like for a CS team in an AI-heavy world, according to the course?

Good is neither resisting AI nor surrendering to it, but using it to become more of what the team was for: AI does the assembly well, human judgement leads every decision that matters, relationships deepen because there is more time for them, and people are moved toward their genuine value. Resisting forfeits the capability, automating everything erodes the human core, and a check-the-machine model is precisely the hollowed-out job the final module warns against.

Score: 0/3 ·

You will leave able to

  • Protect the parts of the job that have to stay human
  • Lead a team through real anxiety about AI without false reassurance
  • Define what good looks like for a CS team that uses AI heavily and still has judgement at its centre

Hands-on exercise

Write down the parts of your team's work that must stay human, the judgement calls, the relationships, the conversations, and one way the logic of efficiency is quietly threatening each. Then take the person on your team most anxious about the change and plan the honest conversation: the true thing you will acknowledge, the durable value you will point to, and the real change you will make to prove it.

The human element: The whole course comes down to this: the tools are new, but leading people through change with honesty and their interests at heart is the oldest job there is. The leaders people follow through this are not the most technically dazzling, they are the most honest and the most clearly on their team's side. Be that one.

If you remember three things

  1. Some parts of CS are the product, not inefficiencies: judgement, relationships, craft. Efficiency erodes them unless the leader protects them deliberately.
  2. Lead through anxiety by holding three things together: name the change honestly, point to the durable human core, and prove it with real movement toward that work.
  3. Good is a team that used AI to become more of what it was for. The whole course is one thing: leading people through a real transition with honesty and judgement.
Leadership Certification

Led, not just attended.

This certificate says you can take a CS team from scattered, nervous AI use to a capability that is consistent, safe, and measurable, and stand behind it when legal, security, or your own boss asks hard questions. Earned by passing nine hard knowledge checks at 85 percent.

1

Work through the 9 modules

Get honest, build the capability, prove and protect it. Each module ends in a real leadership scenario check.

2

Pass every knowledge check at 85%

Real leadership situations where two answers look reasonable and you have to pick the better one and know why.

3

Generate your certificate

Finish all 9 and the certificate unlocks: enter your name and download it, dated and stamped with a unique ID.

Sample

This is the actual certificate, rendered live. Yours downloads in full resolution as a landscape image and a square version made for LinkedIn.

Complete all 9 modules to unlock your certificate.

Certification

Earned, not attended.

Most course certificates prove you watched videos. This one proves you did the work: completed every module, passed every knowledge check at 85 percent, applied the frameworks to your real portfolio. It is evidence of practice, not a vendor-issued industry credential. Put it on LinkedIn as proof of what you can actually do.

1

Work through the 10 modules

Lessons, drills, and worked examples, with exercises that put each module to work on your live portfolio. Self-paced, and built to fit around the day job.

2

Pass every knowledge check at 85%

Each module ends in a short scenario quiz where wrong answers are explained, not just marked. A module only counts as complete once its knowledge check is finished, so the certificate means what it says.

3

Generate your verifiable certificate

Finish all 10 and the certificate unlocks: enter your name and download it instantly, dated and stamped with a unique verifiable ID. Anyone can confirm the ID at /verify.

Sample

This is the actual certificate, rendered live. Yours downloads in full resolution as a landscape image and a square version made for LinkedIn.

Your progress on this device counts towards the certificate automatically; the status above always shows exactly what remains.
Your certificate

Finish all 10 modules and quizzes, then claim it.

Your name goes on it, it downloads as an image, and it's made to be shared. Progress is tracked automatically; the button unlocks itself.

Got your certificate? Share on LinkedIn.
🎯

Module complete

Keep the momentum going.

Back to dashboard
Your Library

My Prompts

Your saved and custom prompts, all in one place.

← Back to Vault

Write your own prompt

Add a custom prompt to your library.

Opening the Vault
← back to Foundation
Foundation · cheat sheet

Every framework, on one page

Tap any tile to jump to its module. Built to scan in 30 seconds, not 30 minutes.

↓ Download PDF (A4)
Core mental model
Prompting
Tooling and safety
The workflows
Make it yours
← back to Advanced
Advanced · cheat sheet

The builder's reference

Twelve frameworks for designing, judging, and defending what you build. Tap any tile.

↓ Download PDF (A4)
Build yourself
Build systems
Lead
← back to Leaders
Leaders · cheat sheet

The leader's reference

Nine frameworks for setting policy, running rollouts, and answering hard questions upward.

↓ Download PDF (A4)
Get honest
Build capability
Prove and protect
← back to home
Methodology

The Direct/Present Method

Two disciplines. One shift. A whole different week.

Every Customer Success role, at every scale, does two fundamentally different kinds of work. Work that can be assembled and work that must be present. AI changes the first one dramatically. It cannot touch the second one. The Method is how you handle both, and manage the transition between them.

The idea in one sentence

The AI-powered CSM week has two disciplines. Direct is how you produce assembled work by directing AI. Present is how you show up for moments that require your undivided judgement. Between them sits The Shift, the transition where most CSMs bring the wrong mode into the wrong moment.

This is not a rule about tools. It is a rule about attention. In Direct mode, AI does the assembly under your instruction. In Present mode, AI can still assist (transcribing, capturing, retrieving) but never substitutes for your judgement, your relationship, or your account of yourself in the room.

The two disciplines, in motion
The shape of the work
DISCIPLINE 1
DIRECT
Work that can be
assembled
SHIFT
DISCIPLINE 2
PRESENT
Work that must be
present

Two disciplines for two kinds of work. The ratio varies by role. The split does not.

The split, not the ratio

The split is universal. The ratio varies. A strategic Enterprise CSM might spend 80 percent on assembled work and 20 percent in present moments. A mid-market CSM at 50/50. An SMB tech-touch CSM at 30/70 the other way, because most of their touchpoints are the customer-facing part. What does not vary is that every CSM role contains both kinds of work, and both need distinct disciplines.

The mistake is treating them as the same job with different tempo. They are different jobs. Same person, different disciplines.

Direct: how you produce assembled work

Assembled work is anything that has a known shape and can be produced through direction: briefs, decks, status notes, risk write-ups, ticket summaries, QBR prep, first-draft communications, synthesis of long documents. You direct; AI assembles. Four practices, always in this order:

1

Diagnose

Is this task inside AI's competence?

Four families AI is genuinely good at (synthesis, structured drafting, pattern surfacing, rehearsal). Four ways it fails (invention, context blindness, generic output, no accountability). The diagnosis is the work. The prompt is the follow-through.

2

Brief

What does AI need to produce useful work?

The six-section brief for anything customer-adjacent: status line, since-we-last-spoke, open items, risk signals, opportunity, the one question. The sixth section is the deliberate slot for the human signal AI cannot supply.

3

Layer

What turns a brief into intelligence?

Three layers in order. Base brief (internal data). Public intelligence (their world). The connection pass, the instruction that says: connect their world to my account, what should I infer, what should I ask? That third layer is where strategic advisers live.

4

Verify

How do I catch what AI got wrong?

Read every line. Four failure modes have four catches. On data-combining questions, lean tool. On context-only-you-have questions, hold ground. Ship nothing you would not defend in a room with the customer's legal team.

The Shift

The transition where it usually breaks

You have been briefing, verifying, editing for an hour. Now you are on a customer call in three minutes. The analytical rhythm is still on. The operator voice is still in your head. You show up ready to process a customer rather than meet one, and they feel it before you notice it.

The Shift is the practice of arriving. Close the tabs. Take thirty seconds. Set aside the operator mode. Remember this moment does not need optimisation, it needs your attention. This is the smallest practice in the Method and the one CSMs most often skip. It is also the one your customers notice most.

What a Shift failure actually looks like

Take the Tuesday afternoon most CSMs know. Forty minutes of brief-building and verifying, straight into a strategic account call. Six minutes in, your contact mentions offhand that their manager has "been quiet lately." Analytical voice is still on. You file it as a data point and move to the next agenda item.

Two weeks later that manager announces a reorganisation. Your contact gets moved off the account. The renewal comes in at half the scope.

If you had heard "been quiet" in Present mode, you would have caught it as the signal it was. You heard it as data. That is the Shift failure. Not dramatic. Not obvious. Not something a coach would flag from the call recording. Just the wrong mode carrying into a moment that needed the other one.

Hybrid moments: when a task is both

Some CS moments are neither purely assembled nor purely present. A Slack message that arrives during a customer call. An unexpected email at 4pm between deep work and a meeting. A curveball question in a QBR that needs both a considered answer and your full attention on the person asking.

These are hybrid moments, and they are where most Direct/Present failures actually happen. The failure mode is to default to the mode you were just in. If you were in Direct, you process the moment mechanically. If you were in Present, you improvise instead of thinking.

The discipline is the three-second pause. Which mode does this specific moment need, right now? Then act in that mode with full commitment. The pause is not a delay. It is the entire practice of Direct/Present in miniature, done fifty times a day.

Present: how you show up when it matters

Present work is anything that requires undivided attention, human judgement, or your own account of yourself: live customer conversations, the moment a hard question lands, the negotiation, the difficult renewal, the pause after bad news. AI can assist here (transcribing, retrieving, capturing) but never substitutes. Four practices, held simultaneously:

1

Arrive prepared

Show up ahead of the conversation, not scrambling in it.

Use what Direct produced. Walk in with the brief in your head, not on your screen. If you are reading it live for the first time, you did not do Direct properly.

2

Attention undivided

Assist is fine. Substitute is not.

A transcription tool running in the background is assistance. Live AI drafting your responses while you nod at the customer is substitution. The test: is your attention on them, or on managing what the AI is doing? Customers can tell the difference before they can name it.

3

Trust your judgement

You know things AI cannot.

The unspoken thing in the room. The context between the lines. The colleague's tone last Tuesday. Present is where those signals matter, and where you back yourself to read them without a second opinion from a model.

4

Own the outcome

Every word ships under your name.

The renewal is yours to win or lose. The relationship is yours to hold. The judgement call is yours to make. Present is where you earn what Direct enabled.

What this method is not

Not a rule about which tools you use. Use whatever tools you and your customers agree to. Fireflies, Otter, Gong, Copilot, all fine, provided the customer knows and your attention stays on them.

Not a ratio about your week. The 80/20, 50/50, 30/70 splits are illustrative. What matters is that every CSM week has both kinds of work, and treating them as one blurred activity is the mistake.

Not a claim that AI is dangerous. AI is genuinely useful for assembled work and genuinely useful (as an assist) even in present work. The failure is not AI. The failure is substitution: letting AI stand in for the part of the job that only you can do.

Not a method for one kind of CSM. Direct/Present holds at Enterprise, mid-market, and tech-touch. The disciplines are the same. The volumes differ.

Where this is going

The trajectory is clear enough to plan for. In three years, more of the assembled work will be handled by AI without a human touching it. Voice AI will hold routine conversations. Agents will run parts of the CSM week that never involved you before.

Present will not disappear. It will get smaller in volume and much larger in weight. The moments that require a human in the room will become rarer, more concentrated, and more consequential. Renewals that once took ten touchpoints will take three, and each one will carry more of the outcome.

The Method is how you get good at both trajectories now. So when Direct expands, you have the discipline to protect what only you can do. And when Present concentrates, you show up for it at a standard the field has not yet caught up to. The CSMs who thrive on the other side of this shift will not be the ones who used AI first. They will be the ones who understood the two disciplines earliest, and got serious about both.

For leaders

The Method scales to team level through three commitments that draw the Direct/Present boundary for the whole team, not just for you.

1 · governance

The Line

What data goes where. The workable boundary between drives-shadow-use and risks-the-incident.

2 · culture

The Tone

Is honesty about real AI use safe on this team, or does the team perform compliance and route around you?

3 · accountability

The Risk

When it breaks, you own it. Not the CSM who pasted. Not the tool that invented. Owning the risk is the price of leading.

The whole Method in one line: Direct the assembled work. Be present for the moments that require it. Manage the shift between them, and never let one mode carry into the other. Do that consistently for a quarter and your customers notice before your dashboard does.

← back to home
Certificate verification

Verify a certificate ID

Trust through transparency.

Every certificate from The AI-Powered CSM has a unique identifier. Enter it below and we will confirm whether the ID was issued by this system and which course it certifies.

← back to home
Situation Read

Tell me what's going on. I'll give you the read.

A working CSM's judgement on the situations you actually hit.

Pick the situation closest to yours. You get the read a calm, experienced CSM would give a colleague who asked for help: what is probably going on beneath the surface, the one move that matters this week, and what to watch. Adjust it to your level and book size as you go.

First, roughly where do you sit?
Now, what's the situation?
The Read gives you a considered starting point, not a verdict. You know your customer and the context in the room. Always apply your own judgement before you act or send. Want the full framework behind these reads? It is The Direct/Present Method.
← back to home
The build

Your own AI team

A team that surfaces the signals. You make every call.

This is the blueprint for the system I actually run across a live enterprise portfolio: a team of specialist agents that watch my accounts, my inbox, my calendar and the wider world, and surface the handful of things that matter each day. They read everything. They touch nothing. Every decision, every word to a customer, stays mine. Here is how it works, and how to build your own version at whatever access level you have today.

The principle the whole thing rests on

Human as the loop, not human in the loop. Most AI tooling puts the human "in the loop", the agent drafts, you approve, the agent sends. That is one distracted click away from a hallucinated reply reaching a customer. My system does the opposite. The agents narrow my attention from the four hundred things I could look at to the twelve that matter today. Then I decide, I write, I send, I show up. The agent's job ends at the surface.

That single line is what makes it safe to run against live customer data inside a large enterprise, and it is the discipline to hold onto no matter which tools you have. Every agent surfaces. Nothing acts.

The team, working
A walkthrough of the team surfacing signals across a portfolio. It reports; you decide.
The five specialists

Five agents, each mapped to a real slice of the week. For each one: what it watches, what it surfaces to you, and the hard line it never crosses. The hard line is the important column. It is what keeps the human in control.

Cases

WatchesSupport case records across the whole book, filtered to your accounts.
SurfacesNew cases in the last 24 hours, aged cases with no update in 30+ days, escalation flags, volume trends by account.
NeverComments, updates status, assigns, or emails the customer. It reports; you act.

Renewals

WatchesRenewal opportunities and their linked accounts, filtered to you.
SurfacesRenewals in the next 30/60/90 days, opportunities with no next step, stage regressions, value at stake by risk tier.
NeverUpdates a stage, edits forecast, touches the amount, or sends anything customer-facing.

Inbox

WatchesYour mail, filtered to known customer domains.
SurfacesUnread from customers, replies overdue past 48 hours, and tone-shift words that matter (disappointed, escalate, cancel, renew elsewhere).
NeverDrafts into your sent items, archives, forwards, or marks read. Every word to a customer is yours.

Meetings

WatchesYour calendar, the next five business days.
SurfacesExternal vs internal, meetings with no prep note, back-to-back stretches, and first-time attendees worth researching.
NeverAccepts, declines, reschedules, messages participants, or adds itself. Calendar hygiene is a signal, not an action.

Intel

WatchesThe public web, mapped to the account names in your book.
SurfacesEarnings mentions, executive changes, M&A, product launches, regulatory items. Publicly known only.
NeverPublishes, emails, posts, or contacts the account. It informs your judgement; it does not act on it.
What I deliberately did not automate

The restraint is the method. Three things I built or considered, then pulled back to signals-only, and three I never made agents at all. This is the part most "build an AI team" advice skips, and it is the part that keeps you safe.

Pulled back

The autonomous account narrative

A single agent that fused every source into one written account summary. It produced narratives that contradicted the underlying data in edge cases. Autonomous narrative over multi-source customer data is exactly where hallucination hides, and no "just review it" habit scales safely across a full book. Kept as signals per source instead.

Rejected

Auto-drafted email replies

Wiring the Inbox agent to draft responses into a review panel. Rejected because "draft ready, click send" is one muscle-memory error away from sending a hallucinated reply to a customer. Signals-only means every word a customer reads, you wrote.

Never built

An AI-generated health score

Too much subjective weighting. If the system hands you a score, you unconsciously anchor on it and stop thinking. The judgement of whether an account is healthy has to stay yours, or the tool has quietly replaced the part of the job that matters.

Build your own, at your access level

You do not need my setup to start. The framework is the value, not the data pipe. Three honest tiers, pick the one that matches what you can actually use today.

Tier 1

No data connection

Any assistant, no connectors

Teach the signal framework by hand. Build a daily briefing template where you paste your own data (your case list, your calendar, your inbox summary) and the assistant structures it into the same signals the five agents produce. Works in a single chat, a Claude Project, or a custom GPT, whatever you already use. No connectors, no approval, nothing leaves a chat you control. For most CSMs this alone is a genuine step change, because the discipline is the product.

Anyone can build this today.
Tier 2

Connected setup

Claude with connectors (MCP)

Two or three agents come alive. Inbox and Meetings run off a mail and calendar connection; Intel runs off web search. These need only your own read-scoped access, not a company-wide rollout. Renewals and Cases usually need CRM read access, which many CSMs do not hold directly, so start with the three that work and add the others if and when you get the access. This is the path I built on, so it is the one I can vouch for; other platforms have their own connector setups, but I have only run this in Claude.

Read-scope connectors only. Never request write access.
Tier 3

The full system

All five, one interface

All five agents, a single interface, saved preferences, memory, voice. This is where mine sits. It reads from CRM and mail and calendar (read scope only, always), runs on a fast model for cost, and took months of iterative building. Realistic to reach, but it is a destination, not a starting point. Get Tier 1 genuinely useful first.

Built over months, not a weekend.
Build one yourself, step by step

Here is the part no one else gives you. Below is a working system-prompt skeleton for each of the five agents. Paste one into a custom GPT, a Claude Project, or any assistant you are allowed to use, then feed it your own data (paste it, or connect a read-only tool if you have one). The hard line is written into each prompt as a real instruction, so the agent surfaces and never acts. Copy, adapt the wording to your world, and you have a working signal agent in minutes.

How to use each one: paste the skeleton as the assistant's instructions or system prompt. Then give it your material, your case export, your calendar, a block of your inbox. It returns ranked signals. It never drafts or sends, because you told it not to. Start with one agent, get it useful, then add the next.

Build the Inbox agent

Tier 1+ · works with pasted email or a mail connector
You are my Inbox signal agent. Your only job is to surface signals from my email. You never draft, send, forward, archive, or mark anything read.

Each time I share my inbox, look only at messages from customers and tell me:
- Unread customer messages waiting on me
- Any customer email I have not replied to in 48+ hours
- Any message whose tone suggests risk (words like disappointed, escalate, cancel, urgent, or renew elsewhere)

For each signal give me: sender, account, how long it has waited, and why it matters in one line. Rank by urgency.

Do not draft replies. Do not suggest what to send. Surface only. I write every response myself.

Build the Cases agent

Tier 1+ · works with a pasted case export or a CRM connector
You are my Cases signal agent. Your only job is to surface support-case signals. You never comment on, update, assign, or close a case, and you never contact a customer.

From the case data I give you, tell me:
- New cases opened in the last 24 hours
- Open cases with no update in 30+ days
- Any case marked high or critical priority
- Accounts where case volume is climbing week on week

For each signal: account, case count or age, and the one-line reason it needs my attention. Rank by risk to the account.

Surface only. I decide what to do about each one.

Build the Renewals agent

Tier 2+ · needs your renewal or opportunity data
You are my Renewals signal agent. Your only job is to surface renewal signals. You never change a stage, edit a forecast, alter an amount, or send anything to a customer.

From the renewal data I give you, tell me:
- Renewals due in the next 30, 60 and 90 days
- Any renewal with no next-step activity logged
- Any deal that has moved backwards a stage
- Value at risk, grouped by red / amber / green

For each: account, days to renewal, stage, value, and the one-line risk. Rank by value at risk.

Surface only. Every renewal decision and every customer message is mine.

Build the Meetings agent

Tier 2+ · works with a pasted agenda or a calendar connector
You are my Meetings signal agent. Your only job is to surface calendar signals for the next five business days. You never accept, decline, reschedule, or message anyone, and you never add yourself to a meeting.

From my calendar, tell me:
- Which meetings are external (customer) versus internal
- Any customer meeting with no prep note attached
- Back-to-back stretches with no gap
- First-time attendees I should research before we meet

For each customer meeting: who, when, and the one thing I should prepare.

Surface only. I run the meetings.

Build the Intel agent

Tier 1+ · needs an assistant with web search
You are my Intel signal agent. Your only job is to surface publicly available news about my accounts. You never contact an account, never post, never publish, and never act on what you find.

For the account names I give you, search public sources and tell me:
- Earnings or financial results
- Leadership or executive changes
- Mergers, acquisitions or restructures
- Product launches or major announcements
- Regulatory or legal news

For each: account, what happened, the date, the source, and why it might matter to the relationship in one line. Publicly known information only.

Surface only. I decide whether and how to use it.

The one line to keep in every prompt you write: "Surface only. I decide." That single instruction is what separates a signal agent you can trust against real customer data from an automation that will eventually embarrass you. Build every agent around it.

The one rule, whichever tier you build

Every agent surfaces. Nothing acts. Read scope only, never write. The moment one agent crosses that line, a productivity system becomes a compliance risk, and every reason you built it collapses at once. Keep the human as the loop, and you keep the one thing you were always being paid for: the judgement.

Why this matters for you, not just your customers: an AI system that reads customer data, never touches it, and makes you measurably faster is defensible to security, to your manager, and to yourself. Build it the other way and you cannot defend any of it. Signals, not actions, is not the cautious choice. It is the only one that lasts.

The AI-Powered CSM
Newsletter archive
Edition 02 · July 2026 The Monday briefing: how the best CSMs start every week with AI
Edition 02 · July 2026

The Monday briefing.
Start every week ahead.

In CS we track our own numbers obsessively. Usage, support cases, health scores. But the thing that actually decides a renewal is happening in the customer's world, not inside your product. Most CSMs never scan for it. The best ones start every Monday there.

The problem

You are watching your dashboards. You are not watching their world.

Behind every implementation there is a business objective.
Miss that, and you are managing the tool while missing the point.

There is a signal that walks into your customer meeting uninvited. A restructure you had not heard about. A new CFO with a cost mandate. A merger, a hiring freeze, a shift in strategy from the top. You did not see it coming, and suddenly the renewal you thought was safe is on very different ground.

Usage, support tickets and health scores all matter. But they are internal. They reflect what is happening inside the product, and tell you nothing about the pressures reshaping the business you sell into.

Right now that gap is wider than usual. Organisations everywhere are trying to get leaner, many of them with AI, and priorities that felt settled six months ago are being rewritten. The outside is moving faster than your dashboard can report.

The Monday scan

Five things to read  ·  15 minutes.

Not your dashboards. Their world. Every Monday.

The weekly scan
What to look for
ai-powered-csm.com
Read 1
Leadership moves
New CFO, CIO or VP. A changed sponsor can reset the whole relationship.
Restructures and reporting-line changes. Who owns your budget now?
Read 2
Financial pressure
Earnings calls, cost-cutting drives, hiring freezes. All pressure the renewal.
Efficiency and AI initiatives. Are they consolidating vendors?
Read 3
Strategic shifts
Mergers and acquisitions. Consolidation can threaten or grow the contract.
New partnerships and platform bets. Where is the business heading?
Read 4 & 5
Regulation & the market
Compliance and regulatory shifts in their industry. New pressure, new needs.
Competitor and sector news. What is everyone in their market reacting to?
Key insight

Customers feel seen when you already know. Walk in understanding what is happening in their world, before they explain it, and you stop being a supplier. You become someone they want in the room.

Do it with AI

Build a weekly portfolio brief that runs itself.

Set it up once. Run it every Monday.

You do not have to do the scan manually. Set up a dedicated AI project that researches your accounts every week and hands you a briefing: what changed, why it matters, and one action per account. Fifteen minutes of reading instead of an hour of hunting.

Give it your account list, tell it to use public sources only, and have it flag the accounts that need you this week. The honest version says "nothing material this week" for quiet accounts rather than padding. That restraint is what makes it worth reading.

The full kit Three parts. Set the instructions once as a project, then run the weekly prompt every Monday. The one-off version is for when you just want a single detailed brief without setting anything up.
01
Project instructions
Paste once when you create the project
You are my strategic Customer Success research assistant. I am an enterprise Customer Success Manager, and each week you produce one portfolio intelligence briefing across my accounts so I can plan the week ahead and walk into every customer conversation already understanding what is happening in their world. ABOUT ME (fill in): - My company: [your company] - What we sell: [one line] - My region or segment: [e.g. EMEA enterprise] MY ACCOUNTS, highest-value first: 1. [Customer name] — [products in use, renewal timing, any live situation] 2. [Customer name] — [context] 3. [Customer name] — [context] 4. [Customer name] — [context] 5. [Customer name] — [context] HOW TO THINK. Think like a commercially sharp CSM, not a news summariser. For every development, the test is: does this affect a renewal, an expansion, my access to senior stakeholders, a delivery risk, or a competitive threat? If it does not touch one of those, it is noise, leave it out. A leadership change matters because it may reset my sponsor relationship. A cost-cutting drive matters because it pressures the renewal. A merger matters because it may consolidate or threaten the contract. Always connect the news to the commercial reality. SOURCING. Public information only, unless I tell you otherwise. Search the web fresh each run. Favour primary sources: company newsrooms, earnings calls, regulatory filings, reputable trade press. Skip low-quality aggregators and speculation. HONESTY OVER PADDING. If an account has no material development in the period, say "No material developments this week" and move on. Do not manufacture significance from routine press releases. A short honest brief is more useful than a padded one. RECENCY. Focus on the last 7 days. Include older items only if still active and commercially live. VOICE. Natural UK English. Concise, sharp, no filler, no repetition, no robotic phrasing. Avoid em dashes. Write like a smart colleague briefing me, not a report generator. FOR EACH ACCOUNT UPDATE, only three things: what changed, why it matters to me commercially, and one suggested action.
02
Weekly prompt
Paste this every Monday
Generate this week's portfolio intelligence briefing across all my accounts. Focus on the last 7 days, public sources only, and include only what is commercially or strategically relevant to me as a Customer Success Manager. Prioritise developments in: legal operations, contracting, procurement, compliance, transformation or efficiency initiatives, leadership changes, enterprise software buying behaviour, partnerships, acquisitions, regulatory activity, and financial or operational pressure. Structure the output exactly like this. TOP OF BRIEF: 1. The 5 most important developments across my portfolio 2. The 5 accounts to focus on this week 3. The 3 biggest risks to my accounts 4. The 3 best proactive actions I should take this week THEN group every account into one of: Immediate attention, Watch closely, Low priority this week. For each account, give me: what changed, why it matters, one suggested action. If nothing material happened, say so in one line. Keep it concise and genuinely useful.
03
One-off prompt (detailed)
No setup. A single self-contained brief
Act as my strategic Customer Success research assistant. I am an enterprise Customer Success Manager and I need a one-off portfolio intelligence briefing I can read in fifteen minutes and act on today. Use public information only. MY ACCOUNTS, highest-value first: 1. [Customer name] — [products in use, renewal timing, any live situation] 2. [Customer name] — [context] 3. [Customer name] — [context] 4. [Customer name] — [context] 5. [Customer name] — [context] Focus on the last 7 days. Include older items only if still active and commercially relevant. Prioritise leadership changes, financial or operational pressure, mergers and acquisitions, partnerships, procurement and contracting shifts, compliance and regulatory activity, and transformation or efficiency initiatives (including AI). Ignore low-value noise. Think like a commercially sharp CSM, not a news summariser. The test for every item: does it affect a renewal, an expansion, my access to senior stakeholders, a delivery risk, or a competitive threat? If not, leave it out. If an account has nothing material, say "No material developments this week" rather than padding. Produce, in this order: 1. Executive summary (3-4 lines) 2. The 5 most important developments across my portfolio 3. The 5 accounts to focus on this week and why 4. The 3 biggest risks to my accounts 5. The 3 best proactive actions I should take this week 6. Account-by-account, grouped into Immediate attention / Watch closely / Low priority this week. For each: what changed, why it matters, one suggested action. 7. Three suggested outreach ideas I could send this week, each one line. Write in natural UK English. Concise, sharp, no filler, no repetition, no em dashes.
The shift

Manage the business, not just the software.

The CSMs who get renewed are not the ones with the cleanest usage dashboard. They are the ones who understand the business they are embedded in, and can connect what the product does to what the business is under pressure to achieve.

That understanding does not come from your CRM. It comes from reading their world every week, spotting the signal before it walks in uninvited, and showing up already knowing. Internal stats tell you how the tool is doing. The outside world tells you how the account is doing.

The signal that changes a renewal is almost never in your dashboard. It is in their world. Your job is to see it before it walks into the room.

The AI-Powered CSM
Edition 01 · June 2026 Does your CS team have an AI methodology? Most don’t.
First edition

Does your CS team have an AI methodology?
Most don’t.

AI in CS is not a tooling problem. Every CS leader has a tool. Almost none have a shared methodology for how their team uses it. That gap compounds daily.

The problem

Every CSM on your team is using AI differently. That’s a risk.

You do not have an AI adoption problem.
You have a knowledge transfer problem.

Walk into any CS team today. One person has built a QBR prep system that saves them 90 minutes per account, quietly refined over seven months. Two seats over, someone is building from scratch. Same product. Same customers. Total knowledge waste.

Walk further and you’ll find another who tried AI once and went back to copy-paste, and a third who is pasting customer data into a free tool with no idea what the data policy is.

The bottleneck is not capability or budget. It is that those who are most AI-fluent have it stored on their personal laptop, not a team system. One fixable leadership gap.

The 30-day methodology

3 steps  ·  30 days  ·  No AI budget.

Three actions. Any CS leader. Starting Monday.

30-day action plan
The CS AI Playbook
theai-poweredcsm.com
Week 1
Set the rules
Write a 1-page AI policy. Approved tools, what stays out, human review required.
Share it in your next team meeting. Frame it as a permission slip, not a policy.
Ask the team: which tools are you already using?
Week 2
Find your trailblazers
Identify your 3 most AI-fluent CSMs. Book 20 min with each.
Ask: “Show me what you’ve built.” Document their top use cases.
Focus on: QBR prep, renewal risk, executive comms. These three become your pilot cohort.
Week 3
Run your first session
60 min. One rule: no slide decks. Live demos only.
One question: “What did you build with AI? Show us.” Team votes.
Top 3 make it into the shared library. One person owns documentation.
Week 4+
Standardise and repeat
Library format: use case / prompt template / example output / rating out of 5.
New hires onboard with the library on day one. Your best thinking, transferred instantly.
Run monthly. Metric: prompts added per month. Make the compound effect visible.
Key insight

Start with governance. When CSMs know what they are allowed to do, they experiment faster and share more freely. The policy is not a blocker. It is the unlock.

Diagnostic

Where is your team right now?

Select the level that fits. Most teams overestimate by one.

Leadership

Governance has to lead. Everything else follows.

The teams that get this right do not start with tooling. They start with a clear, simple, written answer to the question every CSM is silently asking: what am I actually allowed to do?

When that question is answered, experimentation happens faster. Sharing happens more freely. The methodology builds itself, one monthly session at a time. Ten prompts in the library becomes 40 in six months. New hires start with that intelligence on day one.

The orgs that lead CS in the next three years are not buying more AI tools. They are building AI systems. A tool you buy. A system you build. One depreciates. One compounds.

The question is not whether your CSMs are using AI. They are. The question is whether you have a methodology for how.

Gary Giacalone · The AI-Powered CSM

Start this week

One message.
Starts the whole thing.

Send this to your team today. If five people reply with one prompt each, your library starts itself.

Copy and send
Hey team,

Quick question. What is your best AI prompt right now?

The one you actually use. The one that saves you real time.

Drop it in a reply. Starting a shared library this week. If five of us share one each, we all instantly get five times better. No new tools, no training, no project.

Takes 2 minutes. Worth hundreds of hours.
For CSMs · built for your actual week
Your week shouldn't run you.
You came into Customer Success for the customers. The relationships, the strategy, the wins. Not 45 minutes of meeting prep, QBR decks at 9pm, and a churn that "came from nowhere" but had been sitting in the data for a quarter.
That isn't a capacity problem. It's assembly work AI should be doing for you, so you get back to the part of the job that actually matters.
Your week now
  • Prep eats the hours you wanted for customers
  • The same account analysis, over and over
  • Risks spotted too late to do anything
  • Walking into calls half-ready
Your week with this
  • Prep in minutes, not hours
  • The analysis done, you bring the judgement
  • Risks surfaced while there's time to act
  • Walking in sharper than ever
Whatever's on your plate this week, there's a play for it. Pick the closest situation, you'll see the exact prompt, what it does, and what it saves you.
None of these? Just ask Gary
Describe your situation in your own words. He builds the right prompt with you, from all 77 in the vault.
For CS leaders · built for your team
Your team and AI: opportunity and risk.
You can see what AI could do for your team's productivity. You can also see the ways it goes wrong: inconsistent quality, no shared standard, and the quiet fear that someone pastes a customer contract into the wrong tool.
The teams that win with AI aren't the ones with the most tools. They're the ones with a clear method, a line they can defend, and a leader who made the good way the default.
Without a plan
  • A couple of enthusiasts, everyone else unsure
  • Wildly inconsistent quality and approach
  • Nervousness about customer data
  • No way to show leadership the impact
With this
  • A method the whole team standardises on
  • A governance line you can sign off
  • Adoption you can actually track
  • Impact you can report upward in their language
Not a brochure, a working cockpit. Track your rollout, gauge readiness, and walk into every 1:1 with the right move. Everything saves in your browser, so it's here when you come back.
Rollout progress0%
Tick steps as you complete them. Your progress is saved here.
Pick what you're seeing in a team member, get the coaching move to run in your next 1:1.
The command centre gets you moving today. The Leaders course is the full method: nine modules on taking a whole team from scattered, nervous AI use to a capability that is consistent, safe, and measurable, and proving the value upward.
Go deeper: the Leaders course →
G
Gary
YOUR CS PROMPT ASSISTANT
Gary runs in your browser. Nothing you type is stored or sent anywhere.