Ask AI
Add an AI assistant that answers questions from your docs with citations, using your own LLM key, plus MCP servers for coding agents.
Add an ai section to your config and readers get Ask AI: a chat drawer that answers questions from your guides and API reference, citing the pages it used. The same setup serves an MCP server for your docs and one for each API, so coding agents can search your docs and call your API.
Providers and keys
OpenAI, Anthropic, Google or any OpenAI-compatible model; rate limits.
MCP servers
Connect Claude, Cursor or VS Code to your docs and your API.
llms.txt and page actions
Markdown versions of your docs for any AI tool.
Ask AI needs a server
Ask AI, the docs MCP and the API MCP run on the server that serves your docs. That's where your LLM key lives; it never reaches the browser or the build output.
| Served by | Ask AI and MCP |
|---|---|
Your Nest app with mountOrbitDocs | Yes |
Next server mode (output.mode: 'server', with proxy.ts) | Yes |
| The self-hosted platform | Yes |
| A plain static host | No. The Ask AI button hides itself. |
llms.txt, llms-full.txt and the page actions are static files and buttons. They work everywhere. See llms.txt and page actions.
Add Ask AI
Add the ai section
import { defineConfig } from '@orbitdocs/next/config';
export default defineConfig({
site: { title: 'Acme', url: 'https://docs.acme.com' },
ai: {
provider: 'anthropic',
model: 'claude-sonnet-5-5',
apiKeyEnv: 'ANTHROPIC_API_KEY',
askAi: {
greeting: 'Ask anything about the Acme API.',
suggestions: ['How do I authenticate?', 'How do I paginate results?'],
instructions: 'Answer in the voice of the Acme developer team. Prefer TypeScript in examples.',
},
},
});Set the key on the server
ANTHROPIC_API_KEY=sk-ant-…Set it where the docs are served: your Nest app's environment, your Vercel project, your container.
Build
npx orbitdocs buildorbitdocs extract and build write the AI index to .orbitdocs/ai.json: every guide and every API operation as Markdown, and the MCP tool definitions. It holds no secrets. The build copies it to orbitdocs-ai.json, which the server reads and never serves.
Readers now see Ask AI in the top bar.
What readers get
- Ask AI in the top bar, or ⌘ I / Ctrl I anywhere, opens a drawer on the right.
- Before the first question, the drawer shows your
greetingandsuggestionsas buttons. - Answers stream in as Markdown with code blocks. Citations like 1 link to the page section they come from, and the cited pages are listed under the answer.
- Each answer has a copy button. Stop cancels an answer; New chat clears the conversation.
- The conversation survives page navigation within the browser tab (session storage).
How answers are grounded
Ask AI is retrieval-augmented: the model only sees passages from your docs.
- At build time, guides and operations are split into passages at
##and###headings, then into chunks of about 1,500 characters. - For each question, a keyword search (BM25, with a light stemmer and extra weight on titles and headings) picks the best passages. At most two come from one page, and at most
maxSourcesin total. - If a follow-up finds fewer than two passages ("and in Python?"), the search runs again with the previous question added.
- The passages go to the model, numbered, with instructions to answer only from them, cite them as
[1],[2], never invent endpoints or fields, and say so when the docs don't cover the question. Yourinstructionsare added at the end. - Up to 7 earlier messages are sent as history, each cut to 4,000 characters. The question is cut to 2,000.
There is no vector database and no embeddings call: search runs in memory on the docs server.
Private docs stay private
Passages come only from pages the reader may open, using the same rules as the docs server. A partner never gets an answer built from a staff-only page.
Options
Prop
Type
To keep the MCP servers but hide the chat, set askAi.enabled: false.
HTTP endpoints
Build your own chat UI on the same endpoints. Paths are under the base path.
| Endpoint | What it does |
|---|---|
GET /_ai/config | { enabled, greeting, suggestions }. The Ask AI button calls it to decide whether to show itself. |
POST /_ai/chat | Body { "messages": [{ "role": "user", "content": "How do I paginate?" }] }. The last message must be from the user; the last 12 are used. |
/_ai/chat streams newline-delimited JSON (application/x-ndjson):
{"type":"sources","sources":[{"n":1,"title":"Pagination","url":"/guides/pagination/","heading":"Cursors"}]}
{"type":"text","delta":"Pass the `cursor` from the previous page"}
{"type":"text","delta":" as a query parameter [1]."}
{"type":"done"}An error during the answer arrives as {"type":"error","message":"…"}. Errors before streaming return JSON with a status: 400 (bad messages), 404 (Ask AI off), 405 (not POST), 429 (rate limit).
When the AI provider fails (a wrong key, an outage, a rate limit on its side), readers see "The AI provider couldn't answer (HTTP 401). Try again in a moment." with the provider's status code. The provider's own error text isn't shown, because it can echo parts of the request. Calls are retried twice before the error is sent.
Behind a reverse proxy
Answers stream word by word. The server sends X-Accel-Buffering: no so nginx doesn't buffer them. If your proxy buffers responses another way, turn buffering off for /_ai/chat.

