Shadow Engineering: Your Lawyers Are Shipping Code Now
OpenAI published data in June on how agents are being used across its customers, and buried in the adoption curves is a structural change most leadership teams haven’t registered yet.1 Non-developer use of coding-capable agents grew 137x for individual users and 189x inside organisations between August 2025 and June 2026. Legal, finance and recruiting crossed into majority agent usage by April; lawyers and recruiters now put over 85% of their agent output through these tools. And the detail that matters most: more than a quarter of the agent work done by business function employees was engineering and coding tasks.
A caveat before anything else: this is a vendor publishing usage data about its own platform, so treat the precise multiples with appropriate caution. The direction, though, matches what I’m seeing inside companies, and the direction is the point.
From shadow IT to shadow engineering
For a decade the unsanctioned technology problem was shadow IT: teams buying software nobody approved. Annoying, occasionally expensive, but bounded; the worst case was usually duplicate spend and some data in the wrong place.
What’s happening now is different in kind. Business teams aren’t just buying software; they’re building it. The finance analyst who automates the month-end pack has built a data pipeline. The recruiter with a candidate screening agent has built decision software. The lawyer generating contract review tooling has built something that touches privileged information at scale. None of these people report to the CTO, and none of their output passes through anything resembling a development lifecycle.
I’d call this shadow engineering: software creation happening outside the engineering organisation, invisible to the controls that exist precisely because software goes wrong.
Why the engineering controls exist
Code review, testing, change management and release control aren’t bureaucracy; they’re scar tissue. Every one of those practices exists because organisations learned, expensively, what happens without them. Software that silently produces wrong numbers. Changes that break things nobody connected. Systems only one person understands, who then leaves.
Now consider the finance agent that assembles the board pack. If a developer wrote that pipeline it would be reviewed, tested and version controlled. Because an analyst built it with an agent on a Tuesday, it’s reviewed by nobody, tested against nothing, and lives in one person’s account. The board reads its output all the same. That’s not a hypothetical risk profile; it’s the exact one the engineering controls were invented for, now running without them.
And there’s a quieter cost: key person risk is being rebuilt at speed. Every useful agent that lives in one person’s head and login is a small dependency nobody has priced. The analyst leaves; the month-end pack breaks; nobody knows how it worked.
The wrong response and the right one
The wrong response is a ban, and it will fail for the same reason shadow IT bans failed: the productivity is real. The bottleneck these teams are dissolving, waiting for scarce engineering capacity, was one of the genuinely painful constraints in every scaling business. The lawyers and recruiters in OpenAI’s data aren’t misbehaving; they’re solving their own problems at last. Treat your most capable citizen builders as an asset; they’re showing you where the demand for automation actually is.
The right response is ownership, applied with a light touch. Extend your agent register (what’s running, who owns it, what it costs) to include what business teams have built, not just what they’ve bought. Give builders a paved road: a sanctioned platform, scoped data access, and a named owner per agent, so the safe path is also the easy path. And apply one simple gate: anything whose output feeds a system of record or a board decision gets a second pair of eyes, the same way code does. Not everything needs review; the things that move numbers do.
None of this needs a governance committee. It needs someone to own the question, a register that includes the built as well as the bought, and a proportionate rule about what gets reviewed. In most mid sized companies that’s weeks of work, not months.
The companies that get this right will compound the productivity without collecting the incidents. The ones that ignore it will discover their board pack has been assembled by unreviewed software for a year, and they’ll discover it at the worst possible moment.
If you want to know what’s actually being built in your business and which of it deserves a proper production path, that’s what my AI Reality-to-Production Review does. Details at theimpactcto.com, or message me on LinkedIn.