AI Leadership Weekly · Issue #59 · Tuesday 25 November 2025 · 08:00 GMT
Good morning.
Google introduced Gemini 3. Anthropic followed with Opus 4.5. Between them, OpenAI released a Codex model designed to carry work across longer sessions. The release calendar is busy enough to make any shortlist look old. Rather than chase every announcement, this edition separates what is available, what remains restricted and which changes are worth testing against completed business work.
IN 60 SECONDS
Gemini 3 begins rolling out. Google introduced Gemini 3 Pro in preview; the more specialised Deep Think mode initially had more restricted testing access.
Opus 4.5 launches. Anthropic released Opus 4.5 with API launch pricing of $5 per million input tokens and $25 per million output tokens.
Codex tackles longer assignments. OpenAI released GPT-5.1-Codex-Max, using context compaction to support software tasks that continue beyond a single context window.
CEO / COO / CXO CHECKLIST
CEO: Distinguish a preview, a product rollout and a production-ready internal service.
COO: Measure cost per accepted result, including the effort to review it.
CXO: Set checkpoints and release authority for work that continues across multiple steps.
TOP STORIES
1. Gemini 3 begins rolling out
Google · 18 November 2025
What happened. Google announced Gemini 3 on 18 November, beginning with Gemini 3 Pro in preview across several product and developer routes. Its announcement also introduced Deep Think, with access for safety testers ahead of a later subscriber rollout. The performance comparisons were Google’s reported evaluations.
Why it matters. Our take: The useful leadership question is which currently difficult activity deserves another attempt. It might be reviewing a mixed document pack or preparing a prototype from a clear brief. Keep the experiment tied to the recipient’s needs. A richer-looking output is not automatically a better decision or a service your team can depend on.
What to do. Select a few difficult but representative cases and repeat the existing test with the new option. Record the exact product route and preview status. Ask the reviewer to assess the same evidence and quality requirements. Keep a fallback to the current method while the trial is limited and reversible.
2. Opus 4.5 launches
Anthropic · 24 November 2025
What happened. Anthropic released Claude Opus 4.5 on 24 November through its apps, API and major cloud platforms. Standard API launch pricing was $5 per million input tokens and $25 per million output tokens. The company highlighted coding, tool-using work and everyday document tasks; those descriptions are supplier claims, not a business savings forecast.
Why it matters. Our take: A price comparison should use the output your organisation actually accepts. Ask whether a more capable option avoids a failed attempt or reduces correction work; then measure it rather than assuming it. Equally, do not send a simple task to a more elaborate service just because a new release attracts attention.
What to do. Take one routine task and one difficult task. Compare two approved options on the same inputs. Record usage charges, attempts, elapsed time and reviewer effort. State the currency and pricing basis in the trial record. Make a narrow routing decision for the task that benefits, rather than a blanket switch for everyone.
3. Codex tackles longer assignments
OpenAI · 19 November 2025
What happened. On 19 November, OpenAI introduced GPT-5.1-Codex-Max in Codex, with API access described as coming soon at launch. It uses compaction to support work across multiple context windows. OpenAI also recommended restricted access and review of the agent’s work before deployment.
Why it matters. Our take: Treat a long-running assignment as delegated work with an agreed deliverable. “Keep going until it works” is a weak management instruction. A more useful brief describes what may change, what must remain unchanged, which checks must pass and what evidence the reviewer needs to decide whether the result is ready.
What to do. Choose a bounded maintenance task in a test environment. Give it a time or spend limit, an approved scope and a named reviewer. Require a change summary, test results and an explicit list of unresolved items. Keep approval to merge or release with the person responsible for the service.
SIGNALS FROM THE LAST MONTH
7 November · Kimi K2 Thinking is released. Moonshot AI published Kimi K2 Thinking and its model card, adding an open-weight option for reasoning and tool-using workloads. Source
6 November · Gemini adds File Search. Google introduced File Search in public preview, offering managed retrieval over uploaded files through the Gemini API. Source
4 November · A different way to connect tools. Anthropic published an engineering approach using code execution with MCP-connected tools to reduce the information passed through model context. Source
3 November · AWS supplies OpenAI compute. OpenAI announced a seven-year, $38 billion AWS agreement. Compute delivery was separate from distribution of proprietary models through Bedrock. Source
IN BRIEF
More dated updates from the preceding 30 days.
24 November · Claude expands tool use. Anthropic introduced tool search, programmatic tool calling and tool-use examples for developers building agents with larger collections of tools. Source
14 November · Claude previews structured outputs. Anthropic introduced a public beta for structured outputs on selected models, helping developers request responses in a defined data format. Source
13 November · Anthropic reports misuse of Claude. Anthropic reported disrupting a cyber-espionage operation involving Claude. Its account and attribution were the company’s findings, not an independent investigation. Source
12 November · GPT-5.1 reaches ChatGPT. OpenAI introduced GPT-5.1, describing more conversational responses and changes to how the model allocates reasoning effort. Source
THE 15-MINUTE PLAYBOOK
Create a one-page model-change decision
Minutes 0–4: Fix the question. Name the workflow, the current difficulty and the expected improvement. Be specific about who benefits. Do not use a general statement such as “more intelligent” as the only objective.
Minutes 4–8: Preserve a fair comparison. Choose the same approved cases, instructions and evidence for both options. Include an exception. Write the pass criteria before reviewing the results and record any unavoidable difference in the product setup.
Minutes 8–12: Count the whole task. Include preparation, repeated attempts, human review and correction alongside usage. Record failures separately. Avoid projecting a small trial directly into an organisation-wide savings claim.
Minutes 12–15: Set the release boundary. Decide whether to stay, test further or make a limited change. Name the owner, the review date and the fallback. Keep material changes to data access or permitted actions outside the model-only decision.
The output is a defensible choice about one workflow—not a permanent ranking of suppliers.
DATA WAVE MOMENT
Keep the business problem steady while the technology changes around it. Build a small body of test evidence, learn with the people doing the work and make improvements that can be explained in operational terms.
QUESTION FOR READERS
Which difficult task now deserves a fresh test—and what would count as a genuine improvement?
Brought to you by Data Wave — your AI & Data Team as a Subscription.
