AI Leadership Weekly · Issue #102 · Tuesday 22 September 2026 · 08:00 BST

Good morning.

A more natural conversation is an inviting way to make technology easier to use. It can also make an imperfect answer sound unusually assured.

This week’s news puts the interface and the judgement around it side by side. Google introduces new live-model options. IBM asks employees and HR leaders about the changing shape of work. Anthropic brings an external organisation into model evaluation. The useful connection is not “keep a human somewhere in the process”. It is to decide what that person should understand, question and be able to change.

IN 60 SECONDS

Conversation: Google adds new options for live interaction and extended reasoning. People: IBM’s study records concerns and reported work patterns, not proof that skills have deteriorated. Evaluation: Accenture’s role inside Anthropic has funding and accountability arrangements worth noting. This week, make human judgement a designed part of the workflow rather than a reassuring label.

CEO / COO / CXO CHECKLIST

  • CEO: Ask what capability employees should retain as AI takes on more preparation work.

  • COO: Make checking effort visible and allow time for it in the service design.

  • CXO: Clarify who evaluates the system, who funds that work and who acts on its findings.

TOP STORIES

1. Gemini adds live models for different conversational demands

Google, updated 17 September · 15 September 2026

What happened. Google announced Gemini 3.8 Live and Live Extended Thinking on 15 September, with an update on 17 September. The offerings distinguish responsive live interaction from work needing more deliberation. They are model capabilities for building experiences, not evidence that every voice agent is ready for unsupervised customer decisions.

Why it matters. Our take: the right pace depends on the task. A simple clarification should not feel like a committee meeting; a consequential judgement should not be rushed to preserve a conversational rhythm. When designing a voice service, make room for checking, confirmation and handover. Natural speech is useful when it improves understanding, not when it conceals uncertainty behind a fluent delivery.

What to do. Test a proposed voice experience with a straightforward question, an ambiguous request and a case requiring escalation. Observe whether it clarifies rather than guesses. Keep the limits understandable to the user, and evaluate the resulting decision or completed task rather than conversational polish alone.

2. IBM’s workforce study asks what happens to critical thinking

IBM / Oxford Economics · 21 September 2026

What happened. IBM’s 21 September study, conducted with Oxford Economics, surveyed 1,500 HR leaders and 8,800 employees globally between April and June. Sixty per cent of employees reported concern about skill erosion. That is a measure of concern, not evidence that sixty per cent had actually lost skills.

Why it matters. Our take: the practical issue is how a job changes when preparation becomes easier but validation remains important. A person asked to approve an answer needs enough context and skill to challenge it. If checking work is invisible in the plan, the apparent productivity improvement may rely on effort nobody has counted. Training should cover judgement as well as tool operation.

What to do. Ask a team to describe the checking they perform on AI-assisted work. Identify what they need to know to detect an error and where they need access to the original evidence. Build that work into the trial measurement and the learning plan rather than assuming it is negligible.

3. Accenture will embed evaluators inside Anthropic

Anthropic / Accenture · 18 September 2026

What happened. Anthropic announced an embedded evaluation arrangement with Accenture on 18 September, led by Faculty. Anthropic funds Accenture’s work. The arrangement aims to deepen model testing, but that funding relationship should remain visible when describing the external evaluation role.

Why it matters. Our take: bringing another organisation into testing can broaden the scrutiny, but “external” is not a complete description of independence. Buyers should understand who chooses the tests, who can see the findings and who must act on them. Neither an assessment nor a partnership transfers all responsibility for the model or the customer’s deployment to the evaluator.

What to do. For a material supplier assessment, ask who commissioned it, the scope covered and the limits of the conclusions. Request an account of unresolved findings and how they affect the proposed use. Do not let the presence of an evaluator replace your own acceptance decision.

SIGNALS FROM THE LAST MONTH

3 September: Greater frontier-model capability does not remove the need to define acceptable results. OpenAI · 3 September 2026

2 September: Google’s Flash and restricted Cyber offerings have different access conditions. Google · 2 September 2026

3 September: Specialist forecasting illustrates why evaluation should be tied to the decision being supported. Google DeepMind · 3 September 2026

9 September: The Lightwell collaboration directs attention from vulnerability discovery towards validated repair. IBM / LTM / Red Hat · 9 September 2026

IN BRIEF

Further dated updates, including recent context worth keeping in view.

15 September: Gemini Notebook adds study tools including interactive learning formats. Google · 15 September 2026

17 September: Google Labs expands its CC experiment towards groups through a US family-focused waitlist, not an enterprise rollout. Google Labs · 17 September 2026

17 September: Anthropic introduces a beta verification programme for eligible life-sciences institutions and teams. Anthropic · 17 September 2026

10 September lookback: Gemini’s Windows app offers another entry point, with account and permission limits still relevant. Google · 10 September 2026

THE 15-MINUTE PLAYBOOK

Minutes 0–4: Choose a job in which AI prepares material for a person to approve. Write down what the approval is meant to establish.

Minutes 4–8: Ask the reviewer what they actually check and where they get the evidence. Identify anything they cannot verify without recreating the whole task.

Minutes 8–12: Make one improvement to the review surface: clearer source extracts, explicit uncertainty, a comparison with the original request or a visible escalation route.

Minutes 12–15: Agree how checking time and detected errors will be recorded in the next trial. Protect the reviewer’s ability to reject the result. A human approval box is not meaningful when the person lacks information, time or permission to disagree with what the system has produced.

DATA WAVE MOMENT

An effective AI operating model gives people a useful role, not a decorative one. At Data Wave, we ask which judgement remains with the team and what evidence makes that judgement possible. That can lead to more automation in one step and better human involvement in another. The goal is a stronger overall service, not a predetermined count of manual steps.

QUESTION FOR READERS

When a colleague approves an AI-produced result, what are they genuinely equipped to check?

Brought to you by Data Wave — your AI & Data Team as a Subscription.