AI Leadership Weekly · Issue #90 · Tuesday 30 June 2026 · 08:00 BST

Good morning.

Mistral’s document-reading update, Google’s computer-use announcement and Microsoft’s Observability Agent sit at different stages of an automated process. One reads the evidence, another interacts with software, and the third helps investigate what is happening in a service. The connection for leaders is traceability: can the team follow a fact from the original document through to a confirmed action?

IN 60 SECONDS

Mistral releases OCR 4. Mistral introduced OCR 4, adding richer structural information and paragraph-level locations to extracted document content.

Gemini 3.5 Flash adds computer use. Google announced computer-use capabilities for Gemini 3.5 Flash, extending how the model can interact with software interfaces.

Azure adds an Observability Agent. Microsoft announced general availability of Azure’s Observability Agent, applying agents to the investigation and operation of cloud services.

CEO / COO / CXO CHECKLIST

  • CEO: Preserve the source location for every consequential extracted fact.

  • COO: Record the action actually taken, not only the action the assistant intended.

  • CXO: Test whether a colleague can investigate and reverse an incorrect result.

TOP STORIES

1. Mistral releases OCR 4

Mistral AI · 23 June 2026

What happened. Mistral released OCR 4 on 23 June with bounding boxes, block classification and confidence scores alongside extracted content. It supports API and application-level use, with self-managed deployment available to enterprise customers. The scores indicate the model’s confidence; they are not independent proof that an extracted value is correct.

Why it matters. Our take: A useful document workflow should let the reviewer return to the exact part of the source. That is especially important when a small error changes the next action, such as an amount, date or identifier. Combine extraction with business checks: totals that reconcile, required fields that exist and records that match. Keep the original evidence even when the output looks clean enough to use automatically.

What to do. Choose a small set of difficult documents. Ask reviewers to check both the extracted value and its source location. Define which discrepancies should stop the process rather than disappear into a low-confidence queue.

2. Gemini 3.5 Flash adds computer use

Google · 24 June 2026

What happened. Google announced computer use in Gemini 3.5 Flash on 24 June, enabling interaction with graphical interfaces through its developer and enterprise platforms. The announcement discusses controls and safeguards for tool use. Access to the capability does not remove the need for application permissions, isolation or approval of consequential actions.

Why it matters. Our take: Operating a screen can provide an option where a clean integration is unavailable. It also means the system must distinguish an intended action from a confirmed result. A click on “save” is not the same as evidence that the correct record was updated. Design explicit checks around the final state, and avoid broad permissions simply because the interface is familiar to a human operator.

What to do. Trial one reversible, low-impact action in a safe environment. Capture the starting state, intended change and verified ending state. Include a case where the screen or record differs from the expected path.

3. Azure adds an Observability Agent

Microsoft · 23 June 2026

What happened. Microsoft announced general availability of Azure Copilot Observability Agent on 23 June. Built on Azure Monitor, it brings together signals such as logs, metrics, traces and operational context to support investigation. Diagnostic suggestions still need to be assessed; the announcement is not a guarantee that every incident’s root cause will be identified correctly.

Why it matters. Our take: For the business owner, the useful question is whether a failed case can be explained across the whole process. Was the source wrong, the interpretation wrong, the action blocked or the destination unavailable? Connect technical events to a recognisable business record. Otherwise the team may have plenty of logs but no efficient way to understand what happened to the customer’s request.

What to do. Take one failed or reopened case and reconstruct its journey with the operating team. Identify the missing evidence that would have shortened the investigation. Add that evidence before adding more automated actions.

SIGNALS FROM THE LAST MONTH

11 June · IBM and ServiceNow expand collaboration. The companies announced work connecting enterprise data and workflows, with initial new offerings planned for the second half of 2026. Source

9 June · Gemini expands live translation. Google introduced Gemini Live 3.5 Translate, broadening speech-translation capabilities while retaining product-specific availability conditions. Source

8 June · Claude connects with Apple’s framework. Anthropic announced Claude support for Apple’s Foundation Models framework, providing a route beyond the framework’s on-device models. Source

3 June · Gemma 4 gains a 12B model. Google introduced Gemma 4 12B, adding another option for organisations considering local or self-managed inference. Source

IN BRIEF

More dated updates from the preceding 30 days.

18 June · Claude Code adds interactive artifacts. Anthropic introduced artifacts in Claude Code in beta, allowing users to create and share interactive outputs from their development work. Source

16 June · Apptio previews conversational spend analysis. Apptio announced Conversational Insights and other AI capabilities for technology-spend analysis, with the conversational experience initially in preview. Source

16 June · Microsoft sets out its AI approach. Microsoft published its perspective on model diversity, enterprise control and implementation; it was a strategy article, not a new model release. Source

12 June · GPT-5.2 retires from ChatGPT. OpenAI announced retirement of GPT-5.2 models from ChatGPT. A consumer-product change should not be read as an identical API withdrawal. Source

THE 15-MINUTE PLAYBOOK

Minutes 0–4: Select a completed transaction and identify its original request, source documents and required outcome. Choose one with enough complexity to expose a real handover.

Minutes 4–8: Trace the important facts to their source locations. Check how ambiguity, missing information and corrections were recorded.

Minutes 8–12: Follow the resulting action into the destination system. Look for evidence that the intended change actually happened to the correct record.

Minutes 12–15: Ask how the team would detect and recover an incorrect outcome. Assign the most important missing record or check to an owner. Repeat the exercise before expanding the level of automation.

DATA WAVE MOMENT

A dependable automated process connects evidence to action in a way people can inspect. Start with one transaction, preserve the source, define the permitted change and verify the result. That creates a stronger foundation for growth than adding more agents to a process whose failures nobody can explain.

QUESTION FOR READERS

Could we explain a wrong automated action without asking the same system to reconstruct what it thinks it did?

Brought to you by Data Wave — your AI & Data Team as a Subscription.