AI Leadership Weekly · Issue #86 · Tuesday 2 June 2026 · 08:00 BST

Good morning.

Mistral announced a search toolkit and a move into physics AI; Endava’s account of using Codex adds a delivery example. These are not three versions of the same chatbot story. They address finding evidence, applying specialist methods and making expertise easier to reuse. For leaders, the useful starting point is to identify which of those problems is actually slowing the work.

IN 60 SECONDS

Mistral launches Search Toolkit. Mistral introduced components for building search and retrieval, giving developers more control over how agents find relevant information.

Endava describes Codex adoption. An OpenAI case study described how Endava was using Codex and reusable expertise in its delivery work; outcomes were customer-reported.

Mistral moves into physics AI. Mistral introduced a physics AI initiative, extending its focus from language models towards scientific and engineering applications.

CEO / COO / CXO CHECKLIST

  • CEO: Test whether the system found the right evidence before grading its prose.

  • COO: Record the owner and version of expert instructions used in a workflow.

  • CXO: Define an independent check for any prediction or recommendation that affects real operations.

TOP STORIES

1. Mistral launches Search Toolkit

Mistral AI · 28 May 2026

What happened. Mistral announced Search Toolkit on 28 May, combining document ingestion, different retrieval methods and evaluation capabilities. Its approach lets teams compare search configurations against their own test sets. The tooling supports building a retrieval system; it does not guarantee that a particular collection of documents will produce accurate answers.

Why it matters. Our take: A search problem can be surprisingly specific. The correct answer may depend on a document version, an account identifier or a term that means something different inside your organisation. Ask the team to show the retrieved evidence before the generated response. Then check whether the evidence was sufficient, current and permitted. This makes it easier to diagnose a weak answer without treating every failure as a model problem.

What to do. Collect ten recurring questions with their correct source passages. Include one question whose answer is absent. Test retrieval separately and reward the system for recognising when the collection cannot support an answer.

2. Endava describes Codex adoption

OpenAI / Endava · 28 May 2026

What happened. OpenAI’s 28 May Endava case study describes using Codex across requirements, design and delivery, including capturing senior engineering guidance for other colleagues. The account reports faster work on particular engagements. These are customer anecdotes published by a supplier, not a general or independently verified productivity rate.

Why it matters. Our take: The useful question is what makes an expert’s judgement reusable. A statement such as “follow our architecture standards” is not enough. Include examples, trade-offs and conditions under which the normal pattern should not apply. Then let less experienced colleagues challenge the guidance rather than treating it as unquestionable authority. The expert should remain responsible for maintaining the material as the organisation learns.

What to do. Ask a specialist to explain one recurring decision using a good example, a bad example and an exception. Turn that into a short instruction pack and have another colleague apply it to an unfamiliar case.

3. Mistral moves into physics AI

Mistral AI · 27 May 2026

What happened. Mistral’s 27 May physics AI announcement described models that learn from simulation or measurement data to predict physical behaviour. It explicitly distinguishes these models from LLMs and from a complete replacement for traditional solvers. The proposed acceleration of design work still requires verification in the relevant engineering setting.

Why it matters. Our take: Not every valuable AI opportunity is conversational. A business with physical products may benefit from exploring more candidate designs or testing operational scenarios. The crucial question is where an approximation is useful and where precision is non-negotiable. Put that boundary in the hands of the responsible engineering experts, with a clear route from quick screening to validated decisions.

What to do. Identify a design or simulation task where early screening consumes effort. Ask which decisions could use an approximation and which must retain the established verification method. Do not treat a promising prediction as production approval.

SIGNALS FROM THE LAST MONTH

15 May · Databricks tests enterprise agent workloads. Databricks described GPT-5.5 results on enterprise document tasks. The published comparisons were supplier-reported evaluations, not customer-wide guarantees. Source

14 May · IBM proposes integrated AI delivery teams. IBM described small, integrated delivery units intended to connect domain knowledge, engineering and implementation around specific business problems. Source

12 May · IBM adds managed Red Hat inference. IBM announced Red Hat AI Inference and OpenShift Virtualization services on IBM Cloud, expanding enterprise deployment options. Source

7 May · AlphaEvolve reports further applications. Google described additional uses of AlphaEvolve for algorithmic optimisation, extending AI applications beyond conversational assistance. Source

IN BRIEF

More dated updates from the preceding 30 days.

22 May · Mistral introduces remote agents. Mistral added remote agents to Vibe and released Medium 3.5, widening its offer for delegated development work. Source

19 May · Managed Agents adds deployment choices. Anthropic introduced self-hosted sandboxes and MCP tunnels, allowing organisations to separate aspects of execution from managed orchestration. Source

19 May · Google develops a universal shopping cart. Google announced a universal shopping-cart experience, extending its work on AI-assisted discovery and purchasing. Source

19 May · Gemini 3.5 Flash launches. Google introduced Gemini 3.5 Flash while describing Pro as forthcoming, making the distinction between released and planned models important. Source

THE 15-MINUTE PLAYBOOK

Minutes 0–4: Choose one recurring question and write the answer your organisation considers correct. Identify the authoritative source and version.

Minutes 4–8: Run the current process and inspect what it actually retrieved or used. Note missing evidence, irrelevant material and unsupported assumptions separately.

Minutes 8–12: Change one condition: remove the answer, introduce an old version or add a conflicting source. Define the correct response before asking the tool again.

Minutes 12–15: Decide whether the next improvement belongs in source ownership, retrieval, instructions or model capability. Assign it to the appropriate person rather than sending every issue back to the AI developer.

DATA WAVE MOMENT

Useful organisational knowledge is not a folder full of documents or a long prompt. It is evidence that can be found, interpreted and maintained within a real process. Build a small, testable connection between the question, the source and the action, then expand it as the evidence improves.

QUESTION FOR READERS

Which recurring AI mistake begins with the wrong information rather than the wrong reasoning?

Brought to you by Data Wave — your AI & Data Team as a Subscription.