Before your first pilot
Map chunking and metadata for one document type
Share:








Share:




Share:




The bet almost every company is making right now: the last AI pilot stalled because the data was messy, and a cleanup project fixes it.
What’s actually true:
Four distinct pipeline failures cause this, and each one is invisible from a spreadsheet. Here’s what this covers:
| Quick definition. A document doesn’t stay one file once AI touches it. It gets cut into smaller passages, chunks, and each chunk becomes an embedding, a string of numbers capturing what it means, so the system can find it later by meaning, not keyword. The diagram below shows that full path, document, chunks, embeddings and metadata, a retrieval filter, and one of two outcomes depending on whether that filter checked for staleness. |

Chunking decides where one passage ends and the next begins. Most default settings do it by size, roughly every 300 to 500 tokens, with no awareness of what’s inside.
What breaks:
Steps to check this, this week:
Most companies can’t name who’s responsible for a document staying accurate, or when it’s next due for a check.
The five things a file needs before it’s actually AI-usable:
Most companies can name zero of the five for most of what they hold. |
Fix it like this:
This is precisely what we do, before an agent goes anywhere near production data: get every document owned, dated, and tagged so nothing stale slips through unflagged. It’s not a feature we bolt on afterward, it’s the first thing we build. We’ve written before about what happens when an assistant has no memory of how your company actually works, The Amnesiac Genius, and this is the other half of that same problem.
There’s a name for this in production AI: context poisoning. A fluent, confident answer built on a source that’s dead or superseded, delivered in the exact same tone as a correct one.
Why it’s worse than a missing answer:
Here’s the fix:
The diagram below shows one customer record feeding two jobs. A marketing recommendation can run broad and loose, a miss costs nothing. A credit decision needs the identical data tuned for precision, with a mandatory human sign-off. Same data. Different bar.

How to set the bar:
Map chunking and metadata for one document type
Turn on citation grounding for every answer that matters
Tune retrieval per case, gate the expensive ones behind a human
This is exactly what we do. We find the gap before it costs you a decision. Let’s talk.
Share:







We’ve helped teams ship smarter in AI, DevOps, product, and more. Let’s talk.
Actionable insights across AI, DevOps, Product, Security & more