An internal knowledge base is the documentation your staff and your tools answer from: runbooks, policies, how-to guides and decision records. It determines the quality of every AI answer you serve, because a retrieval tool repeats what it finds. A stale knowledge base produces confident wrong answers faster than a human ever could.

TL;DR

  • Your knowledge base is the input to every AI support tool you buy, so its quality sets the ceiling on theirs.
  • One answer per page, a title phrased as the question people actually ask, and the answer in the first two sentences.
  • Duplicate pages are worse than missing ones, because a retrieval tool will pick one and you cannot predict which.
  • Every page needs a named owner and a review date visible on the page itself, not in a separate tracker.
  • If the knowledge base is not current, fix it before buying a chatbot, because deployment will expose every gap at once.

Why this became urgent

For years a stale internal page was a minor annoyance. Somebody followed it, it did not work, they asked a colleague, and the colleague gave them the real answer. That human was silently correcting your documentation every day, and nobody counted it.

Point a retrieval tool at the same pages and that correction disappears. It reads what is written and serves it confidently to everybody. An error that used to reach one person now reaches everyone, with more authority than the colleague had.

This is why knowledge base work has moved from housekeeping to prerequisite. Teams deploying an IT support chatbot find the deployment took an afternoon and the content took six weeks, and the six weeks were the project.

What makes a page answerable

A page written for a human to scan and a page a retrieval system can answer from are different objects. The human brings context, skips headings and infers what applies to them. A retrieval tool pulls a passage and has only that passage.

Four properties decide whether a page works. One answer per page, because a page covering VPN setup for three operating systems returns the wrong one. A title phrased as the question staff ask, so “How do I get on the VPN” rather than “Network Access Policy v4”. The answer in the first two sentences, because the opening is what gets retrieved. And conditions stated in the passage, since “contractors follow a different process” cannot sit two headings above.

Consider a 180-person company whose onboarding page covered laptop setup, VPN, email and expenses in one document with four headings. Staff coped for years. Pointed at a chatbot, expenses questions returned VPN instructions, because the passage boundaries did not match the question boundaries. Splitting it into four pages fixed that without changing a word.

Duplicates are worse than gaps

A missing page produces a clean failure. The tool says it does not know, somebody asks a person, and you learn about a gap. A duplicate page produces an unpredictable answer, and that is considerably worse.

Most companies hold several versions of the same instructions: one in Notion, one in Confluence from a migration, one in a Slack canvas written for a single new starter. All three are plausible. A retrieval tool selects on similarity rather than recency, so you cannot predict which it picks, and the same question can return different answers to two people on one day.

What to do:

  • Search your top twenty question topics and count how many pages answer each one.
  • Pick one canonical page per topic and delete the others rather than archiving them, because archived pages are still indexed.
  • Where a page must stay for history, move it outside the source the tool reads.
  • Decide which single system is the source of truth, and say so publicly to the whole company.
  • Check Slack canvases, Google Docs and personal drives, which is where most duplicates actually live.

Ownership, and making staleness visible

Knowledge bases do not decay because people are careless. They decay because ownership is implied rather than assigned, and because a wrong page looks identical to a right one.

Two mechanisms fix most of it. A named owner on the page itself, not in a spreadsheet, because a separate tracker becomes stale faster than the content it tracks. And a review date displayed on the page, so that a reader can see the thing was last confirmed fourteen months ago and treat it accordingly. Visible staleness is self-correcting in a way that invisible staleness never is.

Then add the one feedback loop that matters. Every question your chatbot failed to answer is a content defect report, and it is the highest quality signal you will get about your own documentation. Review that list weekly. Any question asked twice without a good answer is a page to write, and it comes to you already prioritised by real demand.

Checklist:

  • Named owner visible on every page, with a deputy for anything business critical.
  • Last reviewed date rendered on the page, not stored in a tracker.
  • A review trigger attached to the change itself, so a process change prompts a page check.
  • Weekly review of questions the tool could not answer.
  • Ownership written into somebody’s objectives, because unowned content is optional work.

Where to keep it

The tool matters less than the decision to have exactly one. Notion and Confluence both work and both suffer the same failure, which is that anybody can create a page and nobody deletes one. Guru and Slab are built around verification cycles, which helps directly with the staleness problem. Zendesk Guide and the Freshservice knowledge base keep content next to the ticket queue, which suits teams whose documentation is written by the people answering tickets.

Pricing for this group was not verified for this article, so no figures appear here. Check each vendor’s own page before budgeting.

On the answering side, verified 5 October 2026: Freshservice $19 per agent per month annually with Freddy AI a further $29, Zendesk $55 per agent per month with Copilot a further $50, Crisp free then $45 a month per workspace, and Matram $29 a month flat with unlimited seats. Matram is owned by the same people who publish PeopleOpsHQ. It answers from your documents, which makes it exactly as good as the content above and no better; it does not act on tickets, and has no free tier.

Where to start if it is a mess

Do not audit everything. A full audit of several hundred pages stalls in week two and teaches the company that this work does not finish.

Take the five questions your IT team answers most often and fix only those pages. Split each so it covers one answer, retitle it as the question, move the answer to the top, delete duplicates, add a name and a date. That is a day of work covering a disproportionate share of real volume. Then let the unanswered question list decide what comes next.

Final Thoughts

  • Your knowledge base sets the ceiling on every AI support tool you buy, so treat it as the prerequisite rather than the follow-up.
  • One answer per page, titled as the question, answer in the first two sentences, conditions stated in the passage.
  • Delete duplicates rather than archiving them, and name one system as the source of truth out loud.
  • Put the owner and the last reviewed date on the page itself so staleness is visible to readers.
  • Treat every unanswered question as a prioritised content defect report.
  • Fix the five highest-volume pages first. A full audit will stall and a day of targeted work will not.

Frequently Asked Questions

What is an internal knowledge base?

It is the documentation your own staff answer from: IT runbooks, policies, how-to guides, onboarding instructions and the decision records explaining why things work the way they do. It differs from a customer-facing help centre in audience and tone, and increasingly in function, because it is now also the source that AI support tools read when they answer an employee question. That second role is what changed its importance.

Do we need a knowledge base before buying an AI support tool?

Yes, and this is the most common sequencing mistake in the category. Retrieval tools answer from the documents you point them at, so deploying one against stale or duplicated content produces confident wrong answers at scale rather than better support. Teams that buy first typically find installation took an afternoon and the content work took six weeks, spent with the tool live and underperforming. Fixing the five busiest pages first takes about a day.

What makes documentation work for AI retrieval?

Four structural properties, none of which are about writing more clearly. Each page should answer exactly one question, because a page covering three operating systems returns the wrong one’s instructions. The title should be the question staff actually ask rather than a policy name. The answer belongs in the first two sentences, since the opening passage is what gets retrieved. And conditions, such as a separate contractor process, must sit in the same passage.

Why are duplicate pages a problem?

Because a retrieval tool selects on similarity rather than on recency, which means you cannot predict which version it serves. A missing page fails cleanly and teaches you about a gap, whereas three plausible versions produce different answers to different people on the same day with no indication anything is wrong. Delete rather than archive, because archived pages usually remain indexed.

How do we stop a knowledge base going stale?

Make ownership explicit and staleness visible. Put a named owner on the page itself rather than in a separate tracker, which decays faster than the content it monitors, and render the last reviewed date where readers can see it. Attach the review trigger to the change rather than to the calendar, so that revising a process prompts a check of the pages describing it. Then review the questions your tool failed to answer weekly.

Where should we start if our documentation is a mess?

With the five questions your IT team answers most often, and nothing else. A full audit of several hundred pages reliably stalls within two weeks and teaches everybody that this work does not conclude. Fixing five pages means splitting each so it answers one question, retitling it as the question asked, moving the answer to the top, deleting duplicates and adding an owner and a date. The unanswered question list then decides what comes next.