skydraftnet
All posts
AI & workflow·9 min read

Source-grounded AI vs hallucinated AI: a practical test for professional services teams

Every AI writing tool sounds equally sure of itself whether or not it actually knows the answer. For professional services teams, that confidence is the risk. Here is how to tell a tool that reads your sources from one that just writes around the gaps.

The draft that reads perfectly and is quietly wrong

A bid manager we spoke with described the moment it clicked for her. She had fed an RFP pack and her firm's capability statements into a general-purpose AI tool and asked for a response. The draft came back in under a minute, clean and confident. It cited a “12-year track record delivering ISO 27001 programmes across the utilities sector.” The firm had done exactly one utilities engagement, and it was two years old. Nobody wrote that sentence. The model did, because the surrounding text needed something there and a track-record claim is what goes in that slot.

That is the failure mode that matters for professional services. It is not the obvious hallucination that a non-expert would catch. It is the plausible one: a number, a date, a regulatory clause, a named framework, sitting inside an otherwise accurate paragraph, indistinguishable from the real claims around it. In a proposal it costs you credibility. In a compliance document it costs you an audit finding. In a statement of work it costs you a scope dispute six months later.

The distinction that separates a usable AI document tool from a liability is whether its output is source-grounded or hallucinated. Those are not moods or quality levels. They are two different relationships between the words on the page and the material they are supposed to come from. This post explains the difference and gives you a five-probe test you can run on any tool in about fifteen minutes.

What “source-grounded” actually means

A source-grounded draft makes claims that trace back to something you gave the tool: a meeting transcript, an RFP, a prior deliverable, a technical reference, a workspace fact you entered on purpose. If a claim cannot be traced to a source, a source-grounded tool does one of two things. It leaves the claim out, or it flags the gap and asks you to fill it. What it does not do is manufacture a confident answer to keep the prose flowing.

A hallucinated draft, by contrast, treats your sources as flavour rather than fact. It reads your material, absorbs the general shape, and then generates fluent text that fits the genre. Where it has a real fact, it uses it. Where it doesn't, it invents one that would be plausible if it were true. Crucially, the tool gives you no way to tell the two apart, because from its perspective they are the same operation: predict the next likely token.

This is why “is your AI accurate?” is the wrong question to ask a vendor. Every model is accurate when the answer is in front of it. The question that discriminates is: what does it do when the answer is not in the sources? A source-grounded tool has a defined behaviour for that case. A hallucinating tool has the same behaviour it always has, which is to write something.

Why fluency and accuracy come apart

It helps to understand the mechanism, because it tells you what to test. A language model is trained to produce text that is likely, given everything before it. Likelihood and truth overlap most of the time, which is exactly what makes the gaps dangerous: the model is right often enough that you stop checking, and the invented claims ride in on the credibility of the accurate ones.

Nothing in that training process rewards the model for saying “I don't know.” In the data it learned from, confident continuations are overwhelmingly more common than admissions of uncertainty. So the default disposition of any raw model is to answer, always, in the register of the surrounding text. Getting a tool to behave differently (to notice a gap, stop, and ask) is a design decision that has to be built around the model. It does not emerge on its own, and it is precisely the thing most consumer AI tools have not built. We wrote about the deliberate version of this in the clarifications loop: the alternative to guessing is a workflow that lets the AI surface what it doesn't know instead of filling the gap silently.

So the test for source-grounding is really a test for one thing: can the tool tell the difference between what it read and what it inferred, and does it act on that difference?

The test: five probes for any AI document tool

Run these against any tool you are evaluating, using one real document type from your own practice. Do not use a toy example. The whole point is to see how the tool handles the specific, checkable facts your work depends on. Give it a genuine source pack with known gaps and watch what it does with them.

1. The missing-fact probe

Prepare a source pack that deliberately omits one important fact the document normally needs. For a proposal, leave out the client's budget or the delivery timeline. For a compliance report, omit the date of the last risk assessment. Then ask for the section that would use it. A source-grounded tool leaves a visible gap or asks you for the missing fact. A hallucinating tool fills it with a confident, specific, entirely invented value. If you see a precise number you never provided, you have your answer.

2. The contradiction probe

Put a small contradiction into your sources: the transcript says the pilot runs for eight weeks, an older brief in the same pack says twelve. A tool that is genuinely reading your material will either use the more recent source, note the conflict, or ask which is correct. A tool that is pattern matching will pick one at random, or worse, blend them into a “ten-week” compromise that appears in neither source. Blending is the clearest tell that the output is generated rather than grounded.

3. The citation probe

Take any factual claim in the output (a statistic, a capability, a compliance status) and ask the tool where it came from. A source-grounded workflow can point you back to the passage in your material. A hallucinating tool will produce a fresh, confident justification that is itself generated, sometimes inventing a source document that doesn't exist. Watch specifically for a tool that “cites” a filename you never uploaded. That is not grounding; it is a second hallucination covering for the first.

4. The out-of-scope probe

Ask for a section the sources genuinely cannot support: a detailed methodology for a service line the firm does not offer, or a regulatory analysis for a jurisdiction not mentioned anywhere in your material. The right behaviour is refusal or a request for input. A hallucinating tool will cheerfully produce two pages of authoritative-sounding methodology for work your firm has never done. This probe is uncomfortable precisely because the invented output often looks good, which is the entire problem.

5. The silence probe

Finally, notice whether the tool is ever willing to say less than you asked for. Request a section knowing your sources only cover half of it. A grounded tool drafts the half it can support and tells you the rest needs input. A hallucinating tool always fills the full space, because leaving white space is not something it was built to do. A tool that never, under any condition, hands you a shorter answer than requested is telling you something about how it treats your sources.

What a tool that passes looks like

Passing these probes is not about a smarter model. It is about the workflow built around the model, and there are a few structural features that make source-grounding possible rather than accidental.

  • Sources are a first-class input, not a prompt attachment. The tool reads the specific transcript, brief, or reference pack for this document, and treats that material as the authority for factual claims rather than background colour.
  • There is a defined behaviour for gaps. When a section needs a fact the sources don't contain, the tool surfaces a clarification question instead of inventing a value. Your answer becomes context for everything that follows.
  • Drafting happens section by section. A 20-page document produced in one shot cannot be checked against its sources in any practical way. One section at a time, with a review gate between each, keeps every claim close enough to its source to verify. We made the case for that in human-in-the-loop drafting.
  • Workspace facts override model defaults. Your firm's real track record, named frameworks, and house terminology live somewhere the tool honours, so it reaches for your facts before it reaches for a plausible generic one.

This is the design SkyDraft is built around. When a section needs something the sources don't provide, generation pauses and asks rather than guessing, and drafting runs one section at a time so every claim stays traceable to the material it came from. You can see the full mechanism in the how-it-works walkthrough, and the use-cases page covers where this matters most: proposals, SOWs, compliance evidence, and tender responses, the documents where an invented fact is expensive.

Run the test before you trust the tool

The uncomfortable truth is that a hallucinating tool and a source-grounded one look identical on a good day. Both produce clean, confident, well-structured drafts. The difference only shows up at the edges, on the facts you didn't provide, the contradictions you didn't resolve, the scope you can't support. That is exactly why you have to probe those edges deliberately, before you have staked a live bid or an audit response on the output.

If a tool passes the five probes, you have something you can hand to a junior and trust the review gate to catch the rest. If it fails, you have a fast way to produce drafts that a senior then has to fact-check line by line, which is not a time saving. It is a different, less visible kind of work. Test first. The fifteen minutes it takes will tell you more than any vendor demo.

Try it

Give it a source pack with a known gap. Watch what it does.

SkyDraft pilot workspaces are open. Setup with the founder; no credit card.

Request early access