AI Leadership Weekly · Issue #72 · Tuesday 24 February 2026 · 08:00 GMT

Good morning.

Anthropic and Google refreshed their model offerings this week, while a separate Claude Code Security announcement put vulnerability review in focus. The distinction is useful: general capability improvements and a specialist security workflow should be assessed against different jobs. This edition looks at what was released, what is still in preview and where a practical trial could give a business a meaningful answer.

IN 60 SECONDS

Sonnet 4.6 arrives. Anthropic released Sonnet 4.6, reporting improvements in coding, computer use and reasoning while retaining the Sonnet pricing level.

Gemini 3.1 Pro enters preview. Google introduced Gemini 3.1 Pro in preview, targeting more complex reasoning tasks across its consumer and developer products.

Claude previews code-security review. Anthropic announced a limited research preview of Claude Code Security, designed to identify vulnerabilities and propose fixes for human review.

CEO / COO / CXO CHECKLIST

  • CEO: Retest familiar examples before expanding a model’s authority.

  • COO: Measure the reviewer’s workload as well as the producer’s.

  • CXO: Keep security findings and proposed fixes inside the established remediation process.

TOP STORIES

1. Sonnet 4.6 arrives

Anthropic · 17 February 2026

What happened. Anthropic released Claude Sonnet 4.6 on 17 February, describing improvements in coding, computer use and knowledge work. It also offered a one-million-token context window in beta. Those are launch capabilities and supplier claims; they do not establish that every existing business task will become cheaper, more accurate or suitable for greater autonomy.

Why it matters. Our take: For a document-heavy team, test whether the upgrade makes the important evidence easier to find and use. More input capacity is not a substitute for deciding which version is authoritative or which question needs answering. Keep a reference result that an experienced colleague accepts, so a more polished response cannot quietly change the standard.

What to do. Run a small set of completed assignments through the current and new configurations with the same instructions and source material. Review the outputs without announcing which model produced which. Record missing evidence, factual errors, useful improvements and time needed to reach an acceptable result.

2. Gemini 3.1 Pro enters preview

Google · 19 February 2026

What happened. Google announced Gemini 3.1 Pro on 19 February as an upgrade for complex tasks, with developer preview access and rollout across its product routes. The announcement described progress on reasoning evaluations. It did not mean every account had identical access, or that a benchmark result demonstrated the quality of an organisation’s own decision process.

Why it matters. Our take: Give a reasoning trial a problem that has enough evidence to assess, but is not solved merely by recalling a familiar answer. Ask the reviewer to examine assumptions and alternative explanations. The useful output might be a clearer decision brief or a missing question, not a recommendation accepted because its explanation sounds sophisticated.

What to do. Choose a past decision with a known outcome. Remove knowledge that was unavailable at the time and supply the original evidence. Ask for options, assumptions and unresolved questions. Have the decision owner judge whether the result would have improved their preparation, rather than scoring fluency alone.

3. Claude previews code-security review

Anthropic · 20 February 2026

What happened. On 20 February, Anthropic introduced Claude Code Security in a limited research preview. It described scanning codebases for vulnerabilities and suggesting patches for human review. The announcement was about assistance for defenders; it was not general availability or a claim that generated fixes could safely be applied to production without review.

Why it matters. Our take: Treat a finding as the start of a controlled investigation. A useful service must help the team distinguish a genuine issue from noise, understand the affected code and verify that a fix does not create a different problem. More alerts are not automatically a better outcome when the same people must investigate them.

What to do. Ask the security and engineering owners to agree how preview findings would enter the existing queue. Include a known issue and a benign example in any authorised test. Measure validation effort, false alarms and time to a reviewed fix. Keep deployment approvals separate from the tool’s suggestion.

SIGNALS FROM THE LAST MONTH

5 February · OpenAI introduces Frontier. OpenAI announced an enterprise platform for developing and managing agents, initially working with a limited group of customers. Source

4 February · More agents come to GitHub. GitHub opened a public preview of Claude and Codex coding agents for eligible Copilot users within its existing development environment. Source

2 February · Codex gets a desktop app. OpenAI launched a macOS app for managing parallel coding-agent tasks. Windows availability was not part of the initial release. Source

27 January · Prism puts AI inside scientific writing. OpenAI launched Prism, a free research-writing workspace powered by GPT-5.2 for personal accounts; organisational-plan access was still forthcoming. Source

IN BRIEF

More dated updates from the preceding 30 days.

12 February · Anthropic raises fresh capital. Anthropic announced a $30 billion Series G funding round at a $380 billion post-money valuation. Source

12 February · Codex-Spark targets fast coding. OpenAI introduced GPT-5.3-Codex-Spark as a research preview for Pro users, initially focused on text-based, low-latency coding. Source

12 February · Gemini Deep Think expands. Google offered its updated reasoning mode to AI Ultra subscribers and invited interest in early API access, not general API availability. Source

5 February · Opus 4.6 is released. Anthropic launched Opus 4.6, including a one-million-token context window in beta and updated capabilities for longer assignments. Source

THE 15-MINUTE PLAYBOOK

Refresh the acceptance test, not the whole platform

Minutes 0–4 · Select the evidence. Collect three accepted examples, one difficult exception and one incomplete request. Use material you are authorised to process. Write down the reason each example belongs in the test set before looking at the new outputs.

Minutes 4–8 · Define a useful result. Agree a small rubric: correctness, traceable evidence, appropriate uncertainty and practical usefulness. Add a review-effort measure. Describe a failure that would rule out the new configuration even if it performed well on the other examples.

Minutes 8–12 · Compare without theatre. Keep the brief, source material and reviewer consistent. Present the results without promotional labels. Ask what would have to change before each output could be used, and record the work required rather than simply choosing a favourite.

Minutes 12–15 · Make one bounded decision. Adopt for a defined task, extend the test or retain the current approach. Name the owner and the next review trigger. Do not turn a successful trial of one task into permission for unrelated actions or data access.

DATA WAVE MOMENT

A practical evaluation should help a team decide, not become another research project. Start with work the business recognises, make the acceptance standard visible and keep the decision proportionate to the evidence collected.

QUESTION FOR READERS

Could our review process recognise a genuinely better result—or would it mostly reward a more persuasive presentation?

Brought to you by Data Wave — your AI & Data Team as a Subscription.