AI Leadership Weekly · Issue #84 · Tuesday 19 May 2026 · 08:00 BST

Good morning.

Databricks reported results on enterprise agent work, while IBM addressed both deployment technology and the organisation of delivery teams. The combination is useful because a capable model and a functioning service are different achievements. This week’s test for any proposal is straightforward: what evidence supports the capability claim, and what work remains to make that capability useful in your organisation?

IN 60 SECONDS

Databricks tests enterprise agent workloads. Databricks described GPT-5.5 results on enterprise document tasks.

IBM proposes integrated AI delivery teams. IBM described small, integrated delivery units intended to connect domain knowledge, engineering and implementation around specific business problems.

IBM adds managed Red Hat inference. IBM announced Red Hat AI Inference and OpenShift Virtualization services on IBM Cloud, expanding enterprise deployment options.

CEO / COO / CXO CHECKLIST

  • CEO: Label every performance claim as supplier-reported, trial-measured or observed in operation.

  • COO: Include real exceptions and poor-quality inputs in the evaluation set.

  • CXO: Define who supports the service and who decides that it is ready to expand.

TOP STORIES

1. Databricks tests enterprise agent workloads

OpenAI / Databricks · 15 May 2026

What happened. OpenAI’s 15 May Databricks case study reported GPT-5.5 results on OfficeQA Pro, a benchmark involving difficult enterprise documents. Databricks said the model reduced errors by 46% relative to GPT-5.4 in its agent-harness setting and surpassed 50% accuracy. This is a specific benchmark result, not enterprise-wide accuracy or a 46-percentage-point gain.

Why it matters. Our take: The useful question is whether the benchmark resembles your task. A workflow can fail because the source was parsed incorrectly, the wrong evidence was retrieved or a correct fact was applied to the wrong case. Track those separately. A stronger overall score may still leave a critical failure mode unchanged, and the cost of that failure matters more than an average alone.

What to do. Assemble a small set of your own difficult documents. Record the correct facts and evidence locations before running a comparison. Report errors by type, not just a single success percentage.

2. IBM proposes integrated AI delivery teams

IBM · 14 May 2026

What happened. IBM described Forward Deployed Units on 14 May: teams bringing domain specialists, architects and engineers together with AI-assisted delivery. The announcement argues for continuous execution rather than separate strategy and implementation handovers. Its team-productivity comparisons are supplier claims, not an independently established staffing formula.

Why it matters. Our take: The interesting design choice is keeping the people who understand the problem close to those building the solution. That can make it easier to resolve ambiguity while the software is still changing. But a small team needs access to decisions, data and users; simply reducing headcount will not create those conditions. Define the work the team owns and the support it can call on.

What to do. Identify one delivery handover that repeatedly loses context. Bring the relevant people into a working session around a real example and see whether a decision and a working change can be produced together.

3. IBM adds managed Red Hat inference

IBM · 12 May 2026

What happened. IBM announced Red Hat AI Inference on IBM Cloud on 12 May, alongside a managed virtualisation service. The inference offer is intended to run models as managed resources with identity, logging and other operating capabilities. Those features do not establish that every workload will meet its quality, latency or cost target.

Why it matters. Our take: A managed service can transfer infrastructure work to a supplier, but you still need to understand demand. Is usage steady or bursty? How long can a customer wait? What happens when capacity is unavailable? Compare complete operating options using the same assumptions. A lower apparent setup effort may be valuable, but it should not conceal the long-term service obligations.

What to do. Write a basic demand profile for one proposed application, including peaks and failure handling. Ask suppliers to respond to that profile instead of comparing an assortment of unrelated headline prices.

SIGNALS FROM THE LAST MONTH

28 April · IBM Bob becomes generally available. IBM launched Bob globally as an AI development partner spanning planning, coding, testing, deployment and modernisation. Source

28 April · Microsoft showcases customer deployments. Microsoft published a roundup of customer AI projects, describing how businesses were connecting company information with day-to-day work. Source

23 April · GPT-5.5 launches. OpenAI introduced GPT-5.5, broadening its model offering for professional and coding work; API availability followed the initial announcement. Source

23 April · Managed Agents gains memory. Anthropic added inspectable, persistent memory for Claude Managed Agents, letting developers manage information carried between sessions. Source

IN BRIEF

More dated updates from the preceding 30 days.

7 May · AlphaEvolve reports further applications. Google described additional uses of AlphaEvolve for algorithmic optimisation, extending AI applications beyond conversational assistance. Source

6 May · Uber describes its AI assistant. An OpenAI customer account described Uber’s use of AI to connect assistance with the information and decisions involved in operating its services. Source

5 May · Microsoft studies new working patterns. Microsoft described different patterns of collaboration between employees and agents, focusing on the organisation of work rather than a single product launch. Source

4 May · IBM surveys changing leadership roles. IBM’s study of 2,000 CEOs and equivalent leaders reported changes in AI responsibilities; its results describe that surveyed population. Source

THE 15-MINUTE PLAYBOOK

Minutes 0–4: Choose a proposal and list its three strongest claims. Write the actual source beside each, including who measured it and on what task.

Minutes 4–8: Identify the gap between those tasks and your own. Include data quality, volume, exceptions, permissions and the consequences of an error.

Minutes 8–12: Design the smallest trial that would reduce the most important uncertainty. Define the correct answers independently of the tool being tested.

Minutes 12–15: Agree the next decision and the evidence required for it. A good trial can justify a limited deployment without pretending to settle every question about eventual scale.

DATA WAVE MOMENT

A focused team should shorten the distance between an interesting claim and evidence from real work. Put representative examples, business judgement and delivery expertise together. Keep the test small enough to learn quickly and honest enough that stopping or changing direction remains a useful result.

QUESTION FOR READERS

Which claim in our AI business case has not yet been tested on work that resembles our own?

Brought to you by Data Wave — your AI & Data Team as a Subscription.