skydraftnet
All posts
AI & workflow·10 min read

How to evaluate AI document tools: the questions vendors hope you won't ask

Every AI document tool demos the same way, because every demo is built around the one thing that is easy. Here are twelve questions that tell you whether there is a real drafting workflow underneath, three more about the commercial shape, and what a dodge sounds like.

You have sat through three demos and they were identical

A six-partner engineering consultancy we spoke to earlier this year ran a tool evaluation the way most firms do. The practice lead booked three vendor demos in one week, watched all three, and came out of them unable to tell the products apart. Each demo followed the same arc: paste in a brief, click generate, watch a twelve-page proposal appear in under a minute, scroll through it while the vendor narrates how much time this saves. Each one was genuinely impressive. Each one produced a document nobody in the room would have sent to a client.

That is not the vendors being dishonest. It is the demo format selecting for a single capability. Generating fluent prose fast is the cheapest thing a language model does, and it is the most visually satisfying thing to watch. So that is what gets demoed. Everything that actually determines whether your firm can use the output (whether the draft is grounded in your source material, whether it sounds like you, whether your edits survive, whether the twentieth document is faster than the first) happens outside the forty seconds the demo shows you.

The practical fix is to stop being an audience and start running the demo yourself. Below are the questions worth asking, grouped by what they actually test. Bring your own document to the call. A vendor who cannot let you use your own material in a demo has already answered several of these questions for you.

Four questions about grounding

1. “What happens when my source material doesn't contain the answer?”

This is the single highest-value question in an evaluation, and almost nobody asks it. Take your real brief, delete the section that states the delivery timeline, and ask the tool to draft the timeline section. There are exactly three possible behaviours: it invents a plausible timeline, it writes an evasive paragraph that says nothing, or it stops and asks you what the timeline is.

Only the third is useful. The first is the failure mode that costs firms real money, because an invented date reads exactly like a sourced date and there is no visual signal telling your reviewer which is which. We wrote a longer version of this test in source-grounded AI vs hallucinated AI, including four other probes worth running on the same call.

The dodge to listen for: “our model has very low hallucination rates.” That is a claim about averages. You are asking about a specific behaviour on a specific gap. Make them show you.

2. “Where do the facts in this paragraph come from?”

Pick a sentence in the generated draft that contains a number, a date, or a named commitment. Ask the vendor to show you which piece of your uploaded material that came from. If the answer is a shrug, then every review of every document produced by that tool has to be a full re-verification against the source pack, and you have not saved the time you think you have.

3. “Can I control which sources it reads for this document?”

Real engagements have layered material. A tender response might draw on the tender pack (specific to this bid), your standard methodology document (reusable), and two prior winning proposals (examples of shape, not sources of fact). Those three things should not be treated identically. A tool that dumps everything into one undifferentiated context pile will happily pull a client name from the prior proposal into your new one. Ask how the tool distinguishes material it should quote from material it should only imitate.

4. “What does it do with a contradiction?”

Upload two sources that disagree, which is the normal state of a real source pack. The kickoff notes say eight weeks; the signed variation says ten. Ask for the schedule section. A tool that silently picks one has told you it will silently pick one every time, including on the clause that ends up in front of a client.

Four questions about your firm's standard

5. “Where does my firm's voice live?”

There are two possible answers and they are not close in value. The weak answer is “you can describe your tone in the prompt.” That puts the burden on whoever is drafting, which means the voice is only as consistent as your least experienced user on their busiest day. The strong answer is that voice is a workspace asset: captured once, versioned, applied automatically to every future document without anyone re-typing it.

Ask the follow-up too: how did it get captured? If the answer is a form with a tone dropdown, you are going to get generic output with a formality slider. Extraction from a document you already wrote is a meaningfully different starting point.

6. “Show me the glossary.”

Every firm has house terms. You may call them engagements rather than projects, your discovery stage may be Phase 0, your delivery framework may have a name a partner invented in 2019. A model trained on the open internet will default to the industry-generic word every time unless something overrides it. Ask to see where those overrides are stored, then ask whether they apply to a document generated by a junior who has never heard of the convention. That second half is the real question.

7. “Can I see the template that produced this document?”

A hidden prompt is not a template. What you want to see is a structure you can open and edit: named sections and sub-sections, a guideline per section describing what belongs in it, some notion of which sections are mandatory versus optional, and a checklist of the information the document type requires before drafting starts. That artefact is where your firm's institutional knowledge actually accumulates. If it does not exist, the knowledge stays in the head of whoever is best at prompting, and you have bought a tool rather than built an asset. We make the case for that in the how-it-works walkthrough.

8. “Draft me a fourteen-row compliance matrix.”

This is a deceptively brutal test. List sections (deliverables, compliance matrices, risk registers, requirement responses) are where generic tools fail in three specific ways: they drop items, they invent items you never listed, and they flatten fourteen distinct entries into a paragraph of prose. Give the vendor fourteen numbered requirements from a real tender and count what comes back. Fourteen entries, each addressed on its own terms, is a pass. Anything else is a manual reconciliation job you will do on every bid.

Four questions about the workflow around the draft

9. “What happens when I edit section three and then generate section seven?”

Ask this one out loud and watch the pause. In a one-shot tool, the answer is that section seven never sees your edit, so the correction you made in section three (renaming the workstream, fixing the scope boundary) is contradicted four sections later, and you find it during the partner read at 9pm. In a section-by-section workflow, each section is drafted with the approved versions of the ones before it, so a correction propagates forward instead of getting overwritten. That difference is worth more than any raw speed number on the vendor slide, and we ran the side-by-side in the case for human-in-the-loop drafting.

10. “Where is the review gate?”

There is a real distinction between a tool that produces a whole document and then invites you to read it, and a tool that stops between sections and requires an approval before continuing. Both can be described as “human in the loop” in marketing copy. Only one of them puts the human in the loop while the loop is still running. Ask to be shown the point in the flow where generation halts and waits for a person.

11. “What does the second document cost?”

Every demo shows you document one. Document one is the slowest and least representative, because it includes setup. What you care about is whether document twelve is faster than document one, which only happens if the setup produced something reusable. Ask specifically: after this first document, what exists in the system that makes the next one quicker? If the answer is “you get better at prompting,” the tool has no memory and your firm will not compound anything.

12. “What does the output land in?”

Your deliverable is a Word document with your letterhead, your heading styles, and a table of contents that regenerates correctly. A tool that produces beautiful prose inside its own web editor and exports something you then spend forty minutes restyling has moved the work rather than removed it. Export the demo document, open it in Word, and look at the heading levels and tables before you form an opinion.

Three questions about the commercial shape

These sit outside the product but they decide whether you can deploy it. Ask them of every vendor, in writing, and keep the answers.

  • Who trains on our documents? Not “is it secure,” which every vendor answers yes to. The question is whether your client material becomes training data for anyone, including the underlying model provider, and where the contractual commitment to that is written down. If you work in regulated sectors this question comes before all twelve above.
  • What happens to our templates if we leave? The template, voice, and glossary you build up are your institutional knowledge. Ask whether you can export them and in what format.
  • Who sets this up, and how long do they stay? A tool that requires a well-built template to produce good output and then hands you a blank template screen has outsourced the hard part back to you. Ask what the setup actually involves and who does it in week one.

Running this in a forty-five minute call

You do not need to ask all fifteen. Four of them, run against your own material, will separate the field faster than any feature matrix:

  • The gap test. Delete a fact from your source pack and ask for the section that needs it. Invent, waffle, or ask?
  • The list test. Fourteen requirements in, fourteen entries out?
  • The edit test. Change something in an early section and check whether a later section respects it.
  • The second-document test. What persists after document one that makes document two faster?

A tool that handles all four is doing something structurally different from prose generation. A tool that handles none of them is a chat window with a nicer interface, and you can evaluate that against the chat window your team is already using for free.

Why we published the questions we would be asked

We build SkyDraft for consulting, professional services, and bid teams, so publishing an evaluation checklist is obviously self-interested. Worth saying plainly: we wrote these questions because they are the ones we want to be asked. Every one of them maps to a design decision we made deliberately and can demonstrate on your document, including the ones that make us look slower than a one-shot generator, because there is no “draft the whole thing unsupervised” button and we do not intend to add one.

Use the list on us. Use it on everyone else. If a vendor gets defensive about question one, you have learned the most important thing about the product in the first five minutes. And if you want the reasoning behind our answers before you book anything, the protocol page and the use cases cover the specifics per document type.

Try it

Bring the four tests and your own document. We will run them live.

SkyDraft pilot workspaces are open. Setup with the founder; no credit card.

Request early access