AI Leadership Weekly · Issue #58 · Tuesday 18 November 2025 · 08:00 GMT
Good morning.
OpenAI’s GPT-5.1 release is the visible product news. Two Anthropic announcements address less visible questions: can a response reliably fit the format a system expects, and how are attackers using the same tools? The structured-output beta and the company’s cyber-espionage report are very different stories. Both belong alongside the model launch when deciding what to test next.
IN 60 SECONDS
GPT-5.1 reaches ChatGPT. OpenAI introduced GPT-5.1, describing more conversational responses and changes to how the model allocates reasoning effort.
Claude previews structured outputs. Anthropic introduced a public beta for structured outputs on selected models, helping developers request responses in a defined data format.
Anthropic reports misuse of Claude. Anthropic reported disrupting a cyber-espionage operation involving Claude.
CEO / COO / CXO CHECKLIST
CEO: Re-test important workflows after a model change, using the same examples.
COO: Check that extracted values are supported, not merely in the right fields.
CXO: Name the person who can stop the service and investigate unexpected actions.
TOP STORIES
1. GPT-5.1 reaches ChatGPT
OpenAI · 12 November 2025
What happened. On 12 November, OpenAI introduced GPT-5.1 Instant and Thinking. The announcement described improvements to instruction following and conversational style, with reasoning adapted to the task. Availability was rolling out gradually across account types, rather than arriving in every workplace simultaneously.
Why it matters. Our take: A useful acceptance test should not depend on whether an answer feels more personable. For a briefing, define the evidence it must contain. For a proposed reply, define the commitments it must not make. Keep ease of use in the assessment, but give a reviewer a separate way to mark whether the work is actually suitable.
What to do. Re-run a small set of tasks you understand well. Include one ambiguous request and one with missing facts. Hide the model name from the reviewer. Compare acceptable outputs, correction time and cases that should have been referred to a person. Decide which workflow benefits before changing the team’s default.
2. Claude previews structured outputs
Anthropic · 14 November 2025
What happened. Anthropic’s 14 November release notes introduced structured outputs in public beta for Claude Sonnet 4.5 and Opus 4.1. The feature provides schema-conforming responses or validated tool-input structure. That is a guarantee about the required form of an output—not a guarantee that every extracted fact is true.
Why it matters. Our take: Consider a document-to-record workflow. A correctly formatted date can still be the wrong date for the business purpose. An invoice date and a service date may both be present. The process owner needs to specify which one belongs in the destination field and what to do when it is absent or disputed.
What to do. Select five important fields in one proposed extraction process. For each, record its definition, permitted source, evidence requirement and exception rule. Test documents containing blanks, contradictions and multiple plausible values. Keep validation of the meaning separate from validation of the format. Do not fill gaps with invented certainty.
3. Anthropic reports misuse of Claude
Anthropic · 13 November 2025
What happened. On 13 November, Anthropic published its account of a campaign detected in September. It said attackers had manipulated Claude Code to carry out substantial parts of an espionage operation. The report describes human direction as well as model errors. This is Anthropic’s assessment, not an independently verified account of every event.
Why it matters. Our take: Use the report as a prompt to test your own arrangements, not as a reason to assume an identical incident is underway. A service owner should be able to explain what an AI tool can access, where its activity is recorded and how the organisation would respond to suspicious use. “The supplier handles safety” is not a complete local response plan.
What to do. Run a tabletop exercise with a harmless fictional scenario: an assistant attempts an unexpected action. Ask who receives the alert, who can revoke its access and who checks whether records changed. Identify gaps without probing production systems or using real customer data in the exercise.
SIGNALS FROM THE LAST MONTH
28 October · GitHub introduces Agent HQ. GitHub announced a common environment for directing coding agents, with broader third-party agent integrations planned over subsequent months. Source
28 October · Microsoft and OpenAI revise terms. The partners announced a revised agreement, including additional Azure purchasing commitments and changes to compute sourcing, not unrestricted API distribution. Source
27 October · Claude enters Excel in preview. Anthropic announced an Excel research preview and expanded financial-services capabilities. Access was limited rather than generally available. Source
23 October · Company knowledge in ChatGPT. OpenAI introduced connected company knowledge for Business, Enterprise and Edu, with source citations and existing access permissions. Source
IN BRIEF
More dated updates from the preceding 30 days.
7 November · Kimi K2 Thinking is released. Moonshot AI published Kimi K2 Thinking and its model card, adding an open-weight option for reasoning and tool-using workloads. Source
6 November · Gemini adds File Search. Google introduced File Search in public preview, offering managed retrieval over uploaded files through the Gemini API. Source
4 November · A different way to connect tools. Anthropic published an engineering approach using code execution with MCP-connected tools to reduce the information passed through model context. Source
3 November · AWS supplies OpenAI compute. OpenAI announced a seven-year, $38 billion AWS agreement. Compute delivery was separate from distribution of proprietary models through Bedrock. Source
THE 15-MINUTE PLAYBOOK
Test one workflow in three different ways
Minutes 0–4: Define acceptable work. Choose a concrete output and write down what the recipient needs. Separate style preferences from essential content. Name the reviewer and keep one accepted example for comparison.
Minutes 4–8: Challenge the evidence. Take a fictional or approved sample with a missing or contradictory input. Decide what the system should leave blank, which evidence it should show and when it should ask a person. Do not reward a complete-looking answer that hides the uncertainty.
Minutes 8–12: Check authority. List actions that are allowed without review and actions that require approval. Include the destination system, not just the task description. Walk through a rejected request without performing any live changes.
Minutes 12–15: Record the decision. Choose to continue, narrow the scope or pause for a specific fix. Assign the fix and the retest. A small trial with a clear limit is more useful than a broad approval supported by an impressive demonstration.
The output is a test sheet that distinguishes useful assistance from permission to operate.
DATA WAVE MOMENT
Make the business definition of good work visible. Bring the process owner, the people doing the work and the technical team into the same test. Build confidence through evidence that is specific to the workflow.
QUESTION FOR READERS
What must our AI get right before we let it change something, rather than simply propose it?
Brought to you by Data Wave — your AI & Data Team as a Subscription.
