Forward Deployed Engineer

Your AI pilot didn't fail
because the AI was bad.

It failed because nobody could tell you how often it was wrong. Every company can buy the same AI now — the advantage went to whoever can actually get it working inside a real business, on real work, with real numbers behind it. I'm Robert Drake. That is the entire job I do.

The problem

Almost everyone is failing at this.

If you spent money on an AI project that never made it into daily use, you are not the exception. You are the overwhelming majority.

~95%

of company AI projects never produce a measurable change to the bottom line. That figure comes from an MIT study widely quoted through 2025 and 2026.

People read that number as proof the technology is overhyped. It isn't. The technology works. What fails is everything around it.

Pilots die the same three ways, every time. Nobody mapped how the work actually gets done — only the tidy version in the process document, which is never what happens on a Tuesday. Nobody built a way to measure whether it was right, so “the demo looked good” became the standard for going live. And it was trusted with real decisions on optimism, so the first embarrassing mistake killed the project and the appetite to try again.

None of those are AI problems. They are engineering and judgment problems, and they are the ones I get paid to solve.

The process document is never what actually happens.Which is why I start by watching, not building

How it works

Three steps. In this order. Every time.

The order matters more than anything else here. Most projects skip straight to building, which is exactly why most projects fail.

First, I watch how the work really happens

I sit with the people actually doing the job and follow what they actually do — including the exceptions that live in one person's head and never made it into any document. You end up with a map of the real process and hard numbers on what it costs you today in hours and mistakes.

Sometimes this step ends with me telling you the process isn't worth automating. That is a good outcome. It costs you two weeks instead of two quarters.

Then I prove it works before you trust it

We agree in writing what “correct” means, then collect real cases from your business where the right answer is already known. Now the question stops being “does this seem good?” and becomes “it got 94 of these 100 right, and here are the 6 it missed and why.”

This is the step almost everyone skips. It is the single biggest reason their projects died.

Then it earns its way into your business

Nothing I build starts by touching your live process. It runs quietly alongside your team first, recording what it would have done so we can compare. It only gets more responsibility when the evidence says it deserves it — and every step of the way there is a way to see what it did and undo it.

What you get

You keep everything, including the ability to fire me.

Not a slide deck and a login. Real documentation, written as we go, that lets your own team understand and run what I built after I'm gone.

Here is everything you walk away holding, stage by stage. The plain description is what it does for you; the name beside it is what your technical people will call it.

StageWhat you get
DiscoveryUnderstanding how the work really happens today
  • a map of how the work actually flows, exceptions included (Workflow map)
  • every place it hurts, ranked (Pain-point inventory)
  • what it costs you today, in hours and errors (Baseline metrics)
  • what happens if it goes wrong, and how fast you'd know (Risk register)
ArchitectureDeciding what gets built, and who is allowed to do what
  • a picture of how the pieces connect (System diagram)
  • exactly what information moves between them (Data contracts)
  • what the system may touch, and what it may not (Permission model)
  • how someone could abuse it, and what stops them (Threat model)
EvalsProving it works, with numbers instead of impressions
  • real cases where the right answer is already known (Golden dataset)
  • an agreed written definition of “correct” (Rubrics)
  • every distinct way it breaks, named and counted (Failure taxonomy)
  • what each run costs and how long it takes (Cost & latency report)
DeploymentTurning it on carefully, and being able to turn it off
  • the staged rollout, with what has to be true at each step (Launch plan)
  • alerts when it starts behaving differently (Monitoring)
  • the conditions that shut it off automatically (Rollback triggers)
  • written instructions so your team can run it without me (Runbook)

If a consultant can't hand you this list at the end, you rented a result instead of buying a capability.

Proof

No logos. No testimonials. A working system instead.

Anyone can put client logos on a page. I'd rather show you something running and let you check my work.

The contact form on this site is read by a system I built. It sorts each message, decides how well it fits, and drafts a reply for me — and it does all of that without being allowed to send anything. I read every one.

Then I published its report card, including the 2 tests it fails. Most vendors would have quietly adjusted the test until the score looked better. The whole argument of this site is that you should not trust anyone who does that.

Two ways to work together

Which one are you?