Consider a mid-sized professional services firm: around 60 fee earners, a two-person compliance function, and roughly 40,000 documents across a network share and a document management system — policies, engagement letters, contracts, regulatory correspondence and years of internal guidance. Nothing is missing. Everything is findable in principle. In practice, finding it takes twenty minutes and an interruption to somebody senior.
The scoping week. Rather than start indexing everything, we spend a week on the documents themselves. Which repositories are authoritative, which are archives, how versions are marked, how much of the corpus is scanned rather than text, and which twenty questions people actually ask. The goal is to find the document set worth starting with — not to demonstrate technology.
What tends to surface. A large fraction of the corpus is duplicated or superseded and should not be in scope at all. A meaningful share of older material is scanned at a quality that needs handling before it is usable. And the questions people ask cluster far more tightly than anyone expected — a few dozen recurring shapes, not thousands.
What gets built first. Not everything. One document set — usually current policies and procedures, because it is bounded, the version position is clear, and the compliance team can mark the answers themselves. Their experts write a question set, the assistant answers with citations, and they score whether each answer is grounded in the right passage. If the retrieval is not good enough, we have learned that in three weeks rather than six months.
What production looks like. The assistant is available to fee earners who previously escalated every question. Answers carry citations to the source paragraph. Role-based access means different teams see different document sets, queries are logged, and where confidentiality demands it the whole system runs on the firm's own servers with a locally hosted model.
The point of this scenario is the sequence, not the numbers: understand the corpus first, prove retrieval on one bounded document set, then expand. It is deliberately the opposite of pointing a tool at every file in the organisation and hoping.