AI Leadership Weekly · Issue #60 · Tuesday 2 December 2025 · 08:00 GMT
Good morning.
DeepSeek updated its model family as two less conspicuous announcements tackled the plumbing behind agents. MCP marked its first year with a specification update, and Anthropic described how coding agents can carry progress between sessions. The common thread is continuity: a task needs to survive a new session, a changed connection or an updated model without losing track of the work.
IN 60 SECONDS
How long-running agents hand over. Anthropic described an engineering approach combining environment setup, progress records and incremental work to help coding agents continue across sessions.
MCP marks its first year. The Model Context Protocol published its November specification update, including additions for authorisation and longer-running work.
DeepSeek updates its model line. DeepSeek released V3.2 and a separate Speciale variant.
CEO / COO / CXO CHECKLIST
CEO: Require a visible record of completed and unfinished work.
COO: Give every connected action an owner and a cancellation route.
CXO: Record which model version produced an accepted result.
TOP STORIES
1. How long-running agents hand over
Anthropic · 26 November 2025
What happened. Anthropic published an engineering account on 26 November describing how it improved long-running coding agents. Its approach separates initial environment setup from incremental work and uses progress records and version history to carry state between sessions. These are findings from software-development experiments, not proof that any business process can run unattended.
Why it matters. Our take: Think of the progress record as the handover between shifts. An instruction to finish a whole assignment is not enough when the next session must reconstruct what happened. For a business pilot, require a clear distinction between a completed deliverable, a partial attempt and a blocked task. None should be left to interpretation.
What to do. Choose a task that spans more than one session. Require an end-of-session note containing changes, evidence, open questions and the next safe step. Ask a colleague who did not watch the work to resume it. Their unanswered questions are your improvement backlog.
2. MCP marks its first year
Model Context Protocol · 25 November 2025
What happened. The Model Context Protocol project released its November specification on 25 November. Alongside other changes, it introduced experimental support for tasks that can continue beyond an immediate response, with explicit progress and completion states. A specification release does not mean every existing connector or client already supports those capabilities.
Why it matters. Our take: A connection establishes a way to communicate; it does not settle who can act or what to do when a request stalls. Ask the same questions you would ask about a service desk ticket: was it accepted, who owns it, can it be cancelled, and what proves completion? Keep those questions in the procurement conversation.
What to do. For one proposed integration, ask the delivery team to demonstrate a timeout, cancellation and failed task. Make the user-facing status understandable without technical logs. Check the actual client and server versions before relying on an experimental feature in a promised business service.
3. DeepSeek updates its model line
DeepSeek · 1 December 2025
What happened. DeepSeek’s 1 December changelog states that its deepseek-chat and deepseek-reasoner endpoints were upgraded to DeepSeek-V3.2. It also described a temporary V3.2-Speciale endpoint with a defined December expiry. This is a reminder that an endpoint name, a model release and an enduring service commitment are different things.
Why it matters. Our take: Do not make the person using a business application responsible for noticing a supplier-side change. Put a named owner between release notes and operational decisions. The relevant question is whether your accepted output changes, not whether the new model performs well on someone else’s preferred benchmark.
What to do. List the model aliases used by one live or planned workflow. Record the last test date, who checks release notes and the fallback. Do not build a permanent dependency on a temporary endpoint. Run a few accepted examples again whenever the underlying version changes.
SIGNALS FROM THE LAST MONTH
14 November · Claude previews structured outputs. Anthropic introduced a public beta for structured outputs on selected models, helping developers request responses in a defined data format. Source
13 November · Anthropic reports misuse of Claude. Anthropic reported disrupting a cyber-espionage operation involving Claude. Its account and attribution were the company’s findings, not an independent investigation. Source
12 November · GPT-5.1 reaches ChatGPT. OpenAI introduced GPT-5.1, describing more conversational responses and changes to how the model allocates reasoning effort. Source
7 November · Kimi K2 Thinking is released. Moonshot AI published Kimi K2 Thinking and its model card, adding an open-weight option for reasoning and tool-using workloads. Source
IN BRIEF
More dated updates from the preceding 30 days.
24 November · Opus 4.5 launches. Anthropic released Opus 4.5 with API launch pricing of $5 per million input tokens and $25 per million output tokens. Source
24 November · Claude expands tool use. Anthropic introduced tool search, programmatic tool calling and tool-use examples for developers building agents with larger collections of tools. Source
19 November · Codex tackles longer assignments. OpenAI released GPT-5.1-Codex-Max, using context compaction to support software tasks that continue beyond a single context window. Source
18 November · Gemini 3 begins rolling out. Google introduced Gemini 3 Pro in preview; the more specialised Deep Think mode initially had more restricted testing access. Source
THE 15-MINUTE PLAYBOOK
Write a restartable task brief
Minutes 0–4 · Define the finish. Choose one real assignment and describe the artefact a recipient must receive. Add two observable acceptance checks. Name anything that must remain unchanged, such as the original file or an existing customer commitment.
Minutes 4–8 · Describe the checkpoint. Specify the smallest useful progress record: work attempted, changes made, evidence collected, unresolved questions and next action. Decide where it lives so a colleague does not have to search through an entire conversation.
Minutes 8–12 · Practise an interruption. Stop a trial part-way through and hand the record to another person. Ask them to identify the next safe action. Note missing context rather than coaching them through the gap.
Minutes 12–15 · Assign continuity. Name the owner of the brief, model-change checks and recovery process. Set a limit on retries and expenditure. Agree when work must return to a person instead of continuing indefinitely.
DATA WAVE MOMENT
Make the handover part of the service design. A modest task that can be resumed, checked and explained is a stronger foundation than a longer demonstration nobody else can reproduce.
QUESTION FOR READERS
Could someone who did not watch the AI work safely pick up where it stopped?
Brought to you by Data Wave — your AI & Data Team as a Subscription.
