skydraftnet
All posts
AI & workflow·9 min read

The case for human-in-the-loop AI drafting (no autopilot, by design)

Full-autopilot AI drafting feels faster until you count the review time. Here's why keeping a human in the loop is a design decision, not a limitation, and how to tell the two apart when you're evaluating a tool.

The 2am submission you had to unpick

Here is a scene most bid managers and engagement leads recognise. A tender is due at 5pm. Someone on the team, under pressure, pastes the RFP and last year's response into a chat window, types “write the full response,” and gets back fourteen pages in ninety seconds. It looks finished. Headings, sub-headings, a confident executive summary, a methodology section that name-checks the right frameworks.

Then you read it. The delivery timeline references a phase that doesn't exist in this engagement. The team bios describe two people who left last year. A compliance clause promises a certification the firm doesn't hold. The client's company name is spelled two different ways. None of it is obviously wrong at a glance, which is exactly the problem. You now have to re-read all fourteen pages against the source material, because the one thing you cannot do is trust it. By 4pm you have spent more time verifying the draft than you would have spent writing the sections you actually cared about.

This is the autopilot tax. The document arrived fast and the work moved backwards, from “draft the content” to “audit a plausible stranger's content.” The faster the tool, the more tempting it is to skip the audit. And the audit is the whole job.

What “human in the loop” actually means

The phrase gets used loosely, so it is worth being precise. Human-in-the-loop drafting does not mean a human reads the output at the end. Everyone reads the output at the end. It means the human is a gate during generation: the system produces a bounded unit of work, stops, and waits for a person to accept, edit, or redirect before it produces the next unit.

The distinction matters because of how errors compound. In a one-shot draft, a wrong assumption in section two silently shapes sections three through nine. The AI treats its own earlier output as fact and builds on it. By the time you catch the original error, it has propagated through half the document in ways that are tedious to trace. A review gate after each section stops propagation at the source. You correct the assumption once, and every later section that reads the prior ones as context inherits the correction instead of the mistake.

We wrote about the mechanics of this in the section-by-section vs one-shot comparison. The short version: one-shot is forty seconds to a draft and hours to a deliverable. Section-by-section is slower to the first page and dramatically faster to the last one, because you are never unwinding compounded errors.

“But I'm senior, I can just fix it”

This is the most common objection, and it is half right. A senior consultant genuinely can take a mediocre one-shot draft and knock it into shape, because they are doing the review gate in their head. They read paragraph by paragraph, catch the invented phase, rewrite the timeline, and reconcile the voice. They are the human in the loop. They just are not getting any help holding the loop open.

Two problems follow. First, that senior is now the bottleneck for every document the firm produces, because the correction step lives entirely in their experience and cannot be handed off. Second, it does not transfer. Hand the same one-shot tool to a junior and the review gate collapses, because they do not yet know what a good methodology section for this client looks like. They cannot catch the invented phase because they don't know it's invented. The output ships, and now the firm has a quality problem it can't see.

A workflow that enforces the gate does two useful things at once. It stops the senior from having to hold the whole loop in their head, and it makes the junior's output reviewable in the same structured way. The skill required to operate the system goes down. The floor on output quality goes up. That is the opposite of what autopilot does, which is raise the floor on speed and lower it on trust.

The clarifications loop: the loop working forwards

Keeping a human in the loop is not only about reviewing what the AI wrote. It is also about the AI telling you what it doesn't know before it writes. This is the difference between a tool that guesses plausibly and one that asks.

A concrete example. You are drafting a statement of work and the source brief never specifies who owns intellectual property in the deliverables. A one-shot tool will pick a plausible default, write a confident IP clause, and move on. You may not notice until legal does, or until the client does, which is worse. A system with a clarifications loop stops and asks: “The source material doesn't specify IP ownership. Which applies to this engagement?” Your answer becomes context for that clause and every downstream section that touches it.

This is the loop running forwards instead of backwards. Instead of the human catching errors after the fact, the system surfaces the decisions only a human can make, at the moment they matter. We covered why this is a design choice rather than a limitation of the model in the clarifications loop. Hallucination is what you get when a tool is forced to answer a question it should have asked.

How to tell real human-in-the-loop from theatre

Plenty of tools now advertise “human oversight” and “AI you can trust,” so the phrase alone tells you nothing. When you evaluate a document tool, these are the questions that separate a real review gate from a checkbox on a marketing page.

  • Is there a “draft the whole thing” button that runs unsupervised? If the default path produces a finished document with no stop points, the oversight is optional, which means in practice it will be skipped under deadline. Ask what happens when nobody clicks anything.
  • Can you edit mid-generation and have it stick? A real gate lets you rewrite section three and has section four read your edited version, not the AI's original. If your edits get overwritten or ignored on the next step, the loop is cosmetic.
  • Does it ever refuse to answer? A tool that never says “I don't have enough to write this section” is a tool that invents. The willingness to surface a gap is the clearest signal the loop is real.
  • Where does the source material live? If the system drafts from a transcript, brief, or RFP pack that it cites, you can check its work. If it drafts from “the model's general knowledge,” there is nothing to check against, and review becomes guesswork.
  • Can a junior produce reviewable output? The test of a workflow is not whether your best person can rescue it. It is whether your newest person can produce something a reviewer can sign off without rewriting from scratch.

The trade-off we made on purpose

We built SkyDraft around the gate rather than around speed-to-finish, and it is worth being honest that this is a trade-off. We could ship a faster end-to-end number if we let the model run the whole document unsupervised. We don't. Even in the Generate all path, each section completes before the next begins and you can stop at any step. There is no button that hands you fourteen unverified pages.

The reason is simple. The value was never in faster bad drafts. You can get those for free from any chat window, and the whole firm already has. The value is in faster good drafts, and “good” in professional services has a specific meaning: something a partner will put their name on, a client will sign, or an auditor will accept. None of those outcomes survive a document nobody checked. Keeping the human in the loop is not a concession to caution. It is the only version of “faster” that produces work you can actually send.

Where to start

If your team has been burned by a plausible-but-wrong draft, the instinct is often to swing back to writing everything by hand. That over-corrects. The problem was never the AI writing. It was the AI writing without a gate. Pick one recurring document type, set up a workflow where every section stops for review, and run it against one real engagement. You will notice the difference on the second document, when the corrections you made the first time are already baked into the template. Our use-cases page walks through the document types teams start with.

Try it

A review gate on every section. No autopilot, by design.

SkyDraft pilot workspaces are open. Setup with the founder; no credit card.

Request early access