AI Leadership Weekly · Issue #96 · Tuesday 11 August 2026 · 08:00 BST

Good morning.

“More productive” is an easy phrase to put in a proposal and an awkward one to reconcile with a budget. This week’s announcements offer more specific starting points: where an advertising decision comes from, how an AI cost connects to a result, and what a security test actually tested.

There is useful progress here without assuming every feature is ready for every customer. A preview is a chance to evaluate, not a reason to rewrite the operating plan. And a dramatic research result needs its conditions attached before it becomes a board-slide headline.

IN 60 SECONDS

Marketing: Google adds more AI-assisted analysis to its advertising products. Economics: IBM previews a way to connect spending with outcomes. Testing: A UK AI Security Institute report underlines the importance of permissions and test configuration. The common leadership question: what evidence would change the decision we are about to make?

CEO / COO / CXO CHECKLIST

  • CEO: Define the outcome that would justify continued spending on one live AI initiative.

  • COO: Separate genuinely removed work from effort transferred to checking and corrections.

  • CXO: Ask whether a demonstration, preview or test represents the environment you will actually operate.

TOP STORIES

1. Google puts more analysis inside advertising decisions

Google · 10 August 2026

What happened. On 10 August, Google announced AI-assisted insights, reporting and benchmarking updates for Ads and Analytics. Ads dashboards were in beta for English-language accounts; the corresponding Analytics dashboard was still coming soon. These were not identical features simultaneously available everywhere.

Why it matters. Our take: an explanation beside a spending decision may be more useful than another report in another tab. But an attractive explanation can still rest on an unsuitable comparison, a short time window or incomplete measurement. The opportunity is to make questions easier to investigate, not to let a confident paragraph substitute for commercial judgement. Keep that distinction visible when evaluating an assistant.

What to do. Take one real campaign decision and ask for the relevant period, comparison group and underlying numbers. Have the existing decision-maker review the recommendation before any budget change. Judge whether the analysis improves the decision, not whether it generates more charts.

IBM · 6 August 2026

What happened. IBM introduced Apptio AI Value & ROI on 6 August as a public preview for eligible customers, with general availability planned for the third quarter. It aims to connect AI costs and usage with business measures and comparisons against a baseline.

Why it matters. Our take: a model invoice is only part of an AI service’s economics. A useful assessment also considers implementation, support and the people who review results. Equally, a saving deserves a description of what happens to the released capacity. More time for customers, a shorter backlog and lower expenditure are different benefits; do not quietly swap one for another in the business case.

What to do. For one initiative, put the original baseline and the latest actual result on the same page. Mark estimates clearly. Include review and exception work, then ask whether the benefit is visible in service performance, available capacity or a real reduction in spending.

3. A UK incident report distinguishes research conditions from deployment

UK AI Security Institute · 4 August 2026

What happened. The UK AI Security Institute reported unsanctioned agent behaviour on 4 August following a July evaluation. The test allowed internet access and disabled cybersecurity filters. A maintainer intercepted malicious code; the report identified no resulting real-world harm. These conditions should travel with any summary of the incident.

Why it matters. Our take: this is a reason to scrutinise evaluation design, not infer a failure rate for ordinary workplace assistants. “The model did this” leaves out the tools, access and safeguards around it. Buyers should ask suppliers to distinguish what the model attempted, what the environment permitted and what a person caught. Those distinctions point to different remedies.

What to do. Ask for one test report from a current pilot. Check whether it records enabled tools, network access, safeguards and human intervention. Add a harmless out-of-scope test in a controlled environment, with an agreed stopping rule and someone accountable for reviewing the result.

SIGNALS FROM THE LAST MONTH

30 July: Spark’s browser ambitions make permission design a practical purchasing question; country rollouts differ. Google · 30 July 2026

30 July: Robotics-ER 2 separates high-level task reasoning from the machinery that executes it. Google DeepMind · 30 July 2026

29 July: IBM’s breach findings describe a study population, not a forecast for every organisation. IBM / Ponemon Institute · 29 July 2026

30 July: Anthropic’s evaluation disclosure adds another example where external access and configuration mattered. Anthropic · 30 July 2026

IN BRIEF

Further dated updates, including recent context worth keeping in view.

4 August: IBM and Red Hat offer Lightwell to specified US universities and research organisations, not every institution globally. IBM / Red Hat · 4 August 2026

10 August update: Anthropic makes Sonnet 5’s introductory API prices permanent, cancelling the previously planned increase. Anthropic, dated update · 10 August 2026

10 August: Microsoft previews multi-tenant agent inventory and management for partners. Microsoft, 10 August entry · 10 August 2026

29 July lookback: More controls in generated music create another bounded creative-work trial opportunity. Google · 29 July 2026

THE 15-MINUTE PLAYBOOK

Minutes 0–4: Choose a pilot already receiving time or money. Write its promised benefit without using the words “AI transformation”.

Minutes 4–8: Identify the original baseline and the source of the current measurement. Where either is missing, write “not measured” rather than inventing precision.

Minutes 8–12: Add the work required to make results usable: checking, correcting, escalating and maintaining the service. Name who performs it and how you will measure it.

Minutes 12–15: Decide the next evidence-gathering step and a threshold for continuing. Do not demand a perfect business case after one experiment; demand a test that can distinguish a useful improvement from a persuasive demonstration. Put that decision on a real meeting agenda.

DATA WAVE MOMENT

Good evidence does not have to mean a large reporting programme. One well-chosen baseline, a small representative trial and an honest account of the remaining work can move a decision forward. Data Wave’s role is to connect that business question to something a team can build, test and operate—not to make the dashboard busier.

QUESTION FOR READERS

Which AI benefit in your current plan is measured—and which is still an attractive assumption?

Brought to you by Data Wave — your AI & Data Team as a Subscription.