The board or the investor has asked
Increasingly the question is not whether you are doing AI but what it has changed. A list of tools is a poor answer, and everyone in the room knows it.
Most companies now have AI activity. Far fewer have anything in production that moves a number. This is an independent view of which of yours is which, what is unsafe, and what it would actually take to close the gap. From someone who builds these systems, not only reviews them.
UK | Ireland | Europe | North America
The board has asked what the AI strategy is, and the honest answer is a list of pilots.
That is the common position and it is not a failure of effort. Teams have run experiments, bought licences, stood up a chatbot, and watched a demo work beautifully in a room. What has not happened is the harder part: something running in front of customers or inside a critical workflow, with reliability, cost and data boundaries understood well enough that nobody is nervous about it.
Four things usually bring a CEO to this conversation.
Increasingly the question is not whether you are doing AI but what it has changed. A list of tools is a poor answer, and everyone in the room knows it.
Twelve months of activity, several enthusiastic teams, nothing in production. The pattern is consistent enough to be diagnosable.
People are already using models on company data, through browser tabs and personal accounts, and nobody can say what has left the building. This is the one CTOs are most often accountable for and least able to see.
A vendor, an agency or an internal enthusiast is proposing a platform, and there is no independent view in the room about whether it is the right answer or a cost you will be carrying for three years.
A demo and a production system are different engineering problems, and the distance between them is where almost all AI budgets are lost.
A demo needs to work once, in front of people who want it to work, on data somebody chose. A production system needs to work on the ugly document, the ambiguous question and the case nobody anticipated; it needs to be wrong in ways you can detect; it needs to keep data where it is supposed to stay; and it needs to cost something predictable per transaction rather than whatever the model happens to consume.
The failure is rarely the model. It is retrieval that works on the sample and fails on the corpus, no measurement of whether output is actually correct, data isolation that was never designed because the pilot only had one user, and a cost profile nobody modelled until the first real month's bill.
There is no shortage of people who will assess your AI readiness. Very few of them have shipped a production system into a regulated workflow, and it shows in the recommendations.
An example. On a document-heavy legal platform, pure vector search kept missing exact medical and legal terminology; the right document was not in the top hundred results. The fix was hybrid retrieval, keyword search running alongside vector search, after which the relevant documents were all in the top ten. That is not a strategy insight. It is the kind of thing you only know because you have watched a system fail on real data at two in the morning, and it is the difference between a roadmap that works and one that reads well.
What that experience produces on your engagement: an honest view of whether a use case can be made reliable, what the grounding and evaluation approach needs to be, where the data boundaries have to sit, and what it will genuinely cost to run rather than to demonstrate.
The same method as every engagement here: scope and fee agreed before anything starts, structured interviews and evidence review, findings written against that evidence, and a board-ready readout with a prioritised plan.
We agree what is in scope, who I speak to, and what I get to look at; the pilots, the vendor contracts, the data flows, the spend.
Interviews across leadership, the technology function, the teams actually running pilots and the people quietly using AI without telling anyone. Review of what is deployed, what it is grounded in, what it costs, and where the data goes.
What is real, what is theatre, what is unsafe, and what is worth doing next. Governance and data exposure treated as findings rather than as a compliance appendix.
A report the board can act on, with a prioritised 30/60/90 plan and a clear recommendation on build, buy or stop.
Most of the attention goes to customer-facing AI. The larger unmanaged exposure in most companies is internal: engineering teams generating code faster than anyone reviews it, with no agreed boundary about where that is acceptable.
The operating model I use separates the surfaces. High autonomy where the cost of an error is low and detection is fast; tests, user interface work, documentation, scaffolding. Mandatory human ownership where it is not; identity and access, architecture decisions, anything touching contracts or customer data. Alongside it, instrumentation of review load and defect escape, so a leadership team can see whether velocity is real or whether it has simply moved the work downstream.
This is the part of AI most likely to be raised in your next diligence process, and the part fewest companies can describe.
Independent practice work since 2025, alongside twenty five years leading engineering organisations in PE-backed, FTSE 100 and growth-stage companies.
Designed and built a production-grade retrieval system on AWS Bedrock for document-heavy litigation workflows, handling case files into tens of gigabytes and thousands of documents. Hybrid keyword and vector retrieval, tuned chunking with page-level source tracking for citation, calibrated extraction to prevent the system inferring beyond its evidence, and three-tier data isolation protecting privileged material across a multi-tenant design. Customer trials returned positive feedback and the system performed as designed; the venture did not reach commercial traction and was wound up. The engineering held; the market did not arrive.
Reviewed the agent and integration architecture of a platform serving several hundred enterprise clients. Identified hallucination exposure and multi-tenant data isolation risks, and defined the production hardening and grounding architecture across its enterprise integration surfaces.
AI-assisted extraction and transformation, deliberately balancing model capability against deterministic steps in the places where reliability matters more than flexibility.
Creator of an open-source, self-hosted security questionnaire automation tool with provider-agnostic inference across cloud, local and fully air-gapped deployment, with answer traceability throughout.
No reseller agreements, no platform partnerships, no implementation practice waiting at the end of the report. If the recommendation is that two of your five pilots should stop and the other three do not need a platform, that is what the report says.
This work earns its fee when a decision is close: a platform purchase, a board question you cannot answer, a diligence process inside the next year, or a set of pilots that need either investment or a decision to stop. If you are early and simply curious, the honest answer is that a conversation costs nothing and an assessment would be premature.
The plan is only worth what gets done with it. Where a CEO wants senior judgement available while the work runs, the Impact CTO Advisory Plan continues on a monthly basis; where the situation needs someone building or leading the build, I take that work directly. Both are agreed at the readout rather than sold afterwards.
Fixed fee, agreed in advance. Scope and fee are set before any work begins.
If you need an independent view of whether your AI activity is going anywhere, the first step is a conversation about what you have running and what you are being asked to prove.
Book a CEO Clarity Call