This is what an AI/RAG pipeline sees when it indexes this page — the same output served at https://camelmind-docs.vercel.app/api/llms/features/run-rag-check.Back to doc

Rendered doc

Run RAG Check

Run a local, no-API-key RAG readiness check on a doc page and get chunk-level findings before it ever reaches a retrieval pipeline.

Run RAG Check helps you find retrieval problems in a documentation page before those problems reach an AI or RAG pipeline.

It takes the same AI-readable text shown in AI View, splits it into retrieval-sized chunks, generates test questions, and uses a built-in keyword retriever to see whether those questions can find the right content.

You get chunk-level findings and retrieval scores that help you identify problems such as:

  • Content that is difficult to retrieve with likely user questions
  • Chunks that don't make sense when retrieved on their own
  • Headings that don't provide enough context
  • Tables or code blocks that lose meaning when extracted
  • Broken links that can stop an AI agent from following a documentation path
  • Sections that are too long or too fragmented for effective retrieval
Note
Run RAG Check is a local heuristic, not a replacement for testing your production RAG pipeline. It uses a built-in keyword retriever and does not call an LLM or require an API key.

When to use Run RAG Check

Use AI View and Run RAG Check together:

  1. Open AI View to see the exact text an AI pipeline receives.
  2. Fix formatting or content problems you find there.
  3. Run Run RAG Check to see how that content behaves when split into chunks and retrieved.
  4. Use the findings to improve the page.
  5. Run the check again to verify the changes.

Enable Run RAG Check

Run RAG Check is enabled by default during local development.

To configure it explicitly, add a ragCheck block under ai in camelmind.config.ts:

typescript
const config: CamelMindConfig = {
  // ...
  ai: {
    ragCheck: {
      enabled: true,
      localOnly: true,
      roles: [],
      defaultChunkSize: 500,
      defaultChunkOverlap: 80,
      maxGeneratedQuestions: 12,
    },
  },
}

Configuration fields

FieldRequiredDescription
enabledNoEnables or disables the feature. Defaults to true. Set to false to hide the tool from the toolbar and disable its route.
localOnlyNoKeeps the author tool restricted to localhost. Defaults to true. Setting this to false disables the tool rather than making it remotely accessible.
rolesNoRestricts access to specific roles, such as ["editor", "admin"]. Uses the same role-based access pattern as AI View. An empty array allows any authenticated local user. This has no effect when authentication is disabled.
defaultChunkSizeNoDefault value for Chunk size, in characters. Defaults to 500. Valid range: 100–4000.
defaultChunkOverlapNoDefault value for Chunk overlap, in characters. Defaults to 80. Must be smaller than the chunk size.
maxGeneratedQuestionsNoMaximum number of test questions generated for each run. Defaults to 12.
Note
Run RAG Check inherits your existing document access controls, just like AI View. A user who cannot open a page cannot open its RAG Check either.

Requirements

Run RAG Check is available only when all of the following conditions are met:

  • The dev server is running with NODE_ENV=development, or CAMELMIND_AUTHOR_TOOLS=true is set.
  • The request's Host header is localhost, 127.0.0.1, or [::1], on any port.
  • OFFLINE_MODE is set to false.
  • ai.ragCheck.enabled is set to true.
  • If roles contains values, the current session has at least one of those roles.

When these requirements aren't met, Run RAG Check is hidden from the toolbar and its route returns 404. This means a production reader does not receive a permission error or other indication that the author tool exists.

Warning
CAMELMIND_AUTHOR_TOOLS=true is a development-only escape hatch for testing author tools outside NODE_ENV=development. Do not set it in a production Docker image or shared staging environment. It removes the environment check but does not remove the localhost host check.

Run a check

The Run RAG Check screen has three panes:

  • AI-readable text and chunks — See how your page is divided into retrieval chunks.
  • Controls and summary — Configure chunking, run the check, and review RAG readiness scores and test queries.
  • Findings — Review issues and recommended fixes.

Run RAG Check

The left pane displays the AI-readable version of your page—the same content shown in AI View—divided into the retrieval chunks used by the check.

Selections are synchronized across the screen:

  • Click a line or chunk to select it.
  • Click a finding to jump to the affected chunk.
  • Click a test query to see which chunks were retrieved and jump to the relevant content.

Preview different chunking strategies

Change Chunk size or Chunk overlap in the middle pane to see how different settings affect your chunk boundaries.

Note

Line numbers refer to the generated AI-readable text, including the llms.txt directive prefix. They will not match the line numbers in your original MDX source.


Fixing common findings

Finding categoryFix
retrieval-miss / retrieval-noiseAdd the exact terms users are likely to search for; make the heading and opening sentence more specific.
weak-heading / weak-heading-contextRestate the section's subject in the heading and opening line instead of a generic label or pronoun.
over-split-section / long-chunkRebalance chunk size, or restructure the section so a natural break falls at a heading.
context-less-table / context-less-codeAdd a lead-in sentence before the table or code block explaining what it shows.
dead-linkFix or remove the link — an agent following it hits a dead end.
no-outbound-linksAdd at least one link to related docs so a multi-step task doesn't stall on this page.

Categories shared with AI View (long-paragraph, context-less-code, context-less-table, weak-alt-text, positional-reference, opaque-jsx, weak-heading-context) use the same fixes described on the AI View page.

Run RAG check


Run RAG Check across every doc

You can run a site-wide diagnostic scan across every page in your navigation tree to find low-scoring docs in a single pass:

bash
npx tsx scripts/rag-check-report.ts [threshold]

The rag-check-report.ts script runs the same chunking, question generation, retrieval, and scoring pipeline as the interactive tool against every document in nav.yml. It sorts the results by overall score and outputs pages scoring at or below the specified threshold (default 89).

For each flagged document, the report displays the overall score, the two lowest-scoring dimensions, and up to five findings:

text
--- /features/example-doc (Example Doc) — overall 72 — weakest: crossLinkQuality=40, retrievability=58
    [blocking] dead-link: ...
    [warning] weak-heading-context: ...
    ...and 3 more findings
Note
  • Run rag-check-report.ts locally as a diagnostic tool. CamelMind does not wire this script into package.json or CI build workflows.
  • The script skips entries without a source file (such as noDropdown group placeholders) or pages that fail to load, counting them separately in the final summary line.

  • AI View — preview the raw text a RAG pipeline ingests, without running chunking or retrieval.
  • LLMs.txt — the underlying AI-readable output both tools are built on.

What the AI sees

1> For a complete documentation index, see /llms.txt. To read any public page as Markdown, append .md to the URL.
2 
3Run RAG Check helps you find retrieval problems in a documentation page before those problems reach an AI or RAG pipeline.
4 
5It takes the same AI-readable text shown in [AI View](/features/ai-view), splits it into retrieval-sized chunks, generates test questions, and uses a built-in keyword retriever to see whether those questions can find the right content.
6 
7You get chunk-level findings and retrieval scores that help you identify problems such as:
8 
9- Content that is difficult to retrieve with likely user questions
10- Chunks that don't make sense when retrieved on their own
11- Headings that don't provide enough context
12- Tables or code blocks that lose meaning when extracted
13- Broken links that can stop an AI agent from following a documentation path
14- Sections that are too long or too fragmented for effective retrieval
15 
16<Callout type="note"> Run RAG Check is a **local heuristic**, not a replacement for testing your production RAG pipeline. It uses a built-in keyword retriever and does not call an LLM or require an API key. </Callout>
17 
18---
19 
20## When to use Run RAG Check
21 
22Use **AI View** and **Run RAG Check** together:
23 
241. Open **AI View** to see the exact text an AI pipeline receives.
252. Fix formatting or content problems you find there.
263. Run **Run RAG Check** to see how that content behaves when split into chunks and retrieved.
274. Use the findings to improve the page.
285. Run the check again to verify the changes.
29 
30---
31 
32## Enable Run RAG Check
33 
34Run RAG Check is enabled by default during local development.
35 
36To configure it explicitly, add a ragCheck block under `ai` in `camelmind.config.ts`:
37 
38```typescript
39const config: CamelMindConfig = {
40 // ...
41 ai: {
42 ragCheck: {
43 enabled: true,
44 localOnly: true,
45 roles: [],
46 defaultChunkSize: 500,
47 defaultChunkOverlap: 80,
48 maxGeneratedQuestions: 12,
49 },
50 },
51}
52```
53 
54### Configuration fields
55 
56| Field | Required | Description |
57|---|---|---|
58| `enabled` | No | Enables or disables the feature. Defaults to `true`. Set to `false` to hide the tool from the toolbar and disable its route. |
59| `localOnly` | No | Keeps the author tool restricted to localhost. Defaults to true. Setting this to `false` disables the tool rather than making it remotely accessible. |
60| `roles` | No | Restricts access to specific roles, such as ["editor", "admin"]. Uses the same role-based access pattern as AI View. An empty array allows any authenticated local user. This has no effect when authentication is disabled. |
61| `defaultChunkSize` | No | Default value for <abbr title="The maximum amount of text, measured in characters, included in each retrieval chunk. Larger chunks preserve more context; smaller chunks make retrieval more focused.">Chunk size</abbr>, in characters. Defaults to `500`. Valid range: 100–4000. |
62| `defaultChunkOverlap` | No | Default value for <abbr title="The amount of text, measured in characters, repeated between adjacent chunks. Overlap helps preserve context when important information falls near a chunk boundary.">Chunk overlap</abbr>, in characters. Defaults to `80`. Must be smaller than the chunk size.|
63| `maxGeneratedQuestions` | No | Maximum number of test questions generated for each run. Defaults to `12`. |
64 
65<Callout type="note"> Run RAG Check inherits your existing document access controls, just like AI View. A user who cannot open a page cannot open its RAG Check either. </Callout>
66 
67---
68 
69### Requirements
70 
71Run RAG Check is available only when all of the following conditions are met:
72 
73- The dev server is running with `NODE_ENV=development`, or `CAMELMIND_AUTHOR_TOOLS=true` is set.
74- The request's Host header is `localhost`, `127.0.0.1`, or `[::1]`, on any port.
75- `OFFLINE_MODE` is set to `false`.
76- `ai.ragCheck.enabled` is set to `true`.
77- If `roles` contains values, the current session has at least one of those roles.
78 
79When these requirements aren't met, Run RAG Check is hidden from the toolbar and its route returns `404`. This means a production reader does not receive a permission error or other indication that the author tool exists.
80 
81<Callout type="warning"> `CAMELMIND_AUTHOR_TOOLS=true` is a development-only escape hatch for testing author tools outside `NODE_ENV=development`. Do not set it in a production Docker image or shared staging environment. It removes the environment check but does not remove the localhost host check. </Callout>
82 
83---
84 
85## Run a check
86 
87The **Run RAG Check** screen has three panes:
88 
89- **AI-readable text and chunks** — See how your page is divided into retrieval chunks.
90- **Controls and summary** — Configure chunking, run the check, and review RAG readiness scores and test queries.
91- **Findings** — Review issues and recommended fixes.
92 
93![Run RAG Check](/images/run-rag-check-1.png)
94 
95=== "AI-readable text and chunks"
96 
97The left pane displays the AI-readable version of your page—the same content shown in [AI View](/features/ai-view)—divided into the retrieval chunks used by the check.
98 
99### Navigate between results {/* toc:exclude */}
100 
101Selections are synchronized across the screen:
102 
103- Click a **line** or **chunk** to select it.
104- Click a **finding** to jump to the affected chunk.
105- Click a **test query** to see which chunks were retrieved and jump to the relevant content.
106 
107### Preview different chunking strategies {/* toc:exclude */}
108 
109Change **Chunk size** or **Chunk overlap** in the middle pane to see how different settings affect your chunk boundaries.
110 
111<Callout type="note">
112Line numbers refer to the generated AI-readable text, including the `llms.txt` directive prefix. They will not match the line numbers in your original MDX source.
113</Callout>
114 
115=== "Controls and summary"
116 
117Use the middle pane to configure chunking, run the check, and review the results.
118 
119### Configure chunking {/* toc:exclude */}
120 
121| Control | Description |
122| :--- | :--- |
123| **Chunk size** | The maximum amount of text, measured in characters, included in each retrieval chunk. Select **300**, **500**, or **800**, or enter a custom value from **100–4000**. |
124| **Chunk overlap** | The amount of text, measured in characters, repeated between adjacent chunks. Overlap helps preserve context when important information falls near a chunk boundary. Select **0**, **50**, **80**, **100**, or **150**, or enter a custom value smaller than the chunk size. |
125 
126<Callout type="tip">
127Start with the default **500-character chunk size** and **80-character overlap**, then experiment with different values if your content produces retrieval or chunk-independence findings.
128</Callout>
129 
130### Run the check {/* toc:exclude */}
131 
132Click **Run check**. CamelMind performs the following steps locally:
133 
1341. Converts the page into AI-readable text.
1352. Splits the text into retrieval chunks using your selected settings.
1363. Generates representative test questions from the page.
1374. Runs the built-in keyword retriever against those questions.
1385. Evaluates the results and displays RAG readiness scores and findings.
139 
140No API key or external AI service is required.
141 
142### Review the summary {/* toc:exclude */}
143 
144After the check completes, the summary displays a **RAG readiness score from 0–100**, calculated from six dimensions:
145 
146| Score | Weight | What it measures |
147| :--- | :---: | :--- |
148| **Retrievability** | 35% | How often the expected chunk appears in the top three results for generated test questions, including penalties for retrieval noise. |
149| **Chunk independence** | 20% | Whether chunks remain understandable when retrieved without their surrounding content. |
150| **Grounding readiness** | 15% | Whether the page provides enough explicit context for an AI system to answer confidently. This is a heuristic; no answer is actually generated or evaluated. |
151| **Format robustness** | 15% | How well the page preserves meaning when converted into the plain text an AI system receives. |
152| **Query coverage** | 5% | How often the expected chunk appears anywhere in the top five retrieval results. |
153| **Cross-link quality** | 10% | Whether the page connects effectively to other documentation through valid outbound links. |
154 
155If a run contains a **blocking** finding, the overall score is capped at **70**, regardless of the weighted score. This prevents a serious failure—such as a dead link or failed retrieval—from being hidden by strong scores in other areas.
156 
157Hover over the information icon next to a score to see its full description.
158 
159### Review test queries {/* toc:exclude */}
160 
161Below the summary, **Test queries** shows the questions generated from your page.
162 
163Each query is marked:
164 
165- **Pass** — The expected content was retrieved successfully.
166- **Warning** — The expected content was retrieved, but other chunks created retrieval noise.
167- **Fail** — The expected content was not retrieved successfully.
168 
169Click a test query to see which chunks were retrieved and jump to the relevant content in the left pane.
170 
171<Callout type="tip">
172Every score is labeled **Heuristic estimate**. Run RAG Check uses a keyword-based retriever as a local proxy for real retrieval. The results can help you identify potential problems, but they are not a substitute for testing your documentation against your production RAG stack.
173</Callout>
174 
175=== "Findings"
176 
177The right pane lists issues detected during the check and groups them by severity.
178 
179- **Blocking** — An outright failure that can prevent successful retrieval or navigation, such as a dead link or a test question that fails to retrieve its expected chunk.
180- **Warning** — A likely retrieval, context, or grounding problem that should be reviewed.
181- **Note** — A lower-confidence observation that may still be worth addressing.
182 
183Each finding includes:
184 
185- **What went wrong**
186- **A recommended fix**
187- **The affected line or chunk**, when applicable
188 
189Click a finding to jump directly to the relevant chunk in the left pane.
190 
191===
192 
193---
194 
195## Fixing common findings
196 
197| Finding category | Fix |
198|---|---|
199| `retrieval-miss` / `retrieval-noise` | Add the exact terms users are likely to search for; make the heading and opening sentence more specific. |
200| `weak-heading` / `weak-heading-context` | Restate the section's subject in the heading and opening line instead of a generic label or pronoun. |
201| `over-split-section` / `long-chunk` | Rebalance chunk size, or restructure the section so a natural break falls at a heading. |
202| `context-less-table` / `context-less-code` | Add a lead-in sentence before the table or code block explaining what it shows. |
203| `dead-link` | Fix or remove the link — an agent following it hits a dead end. |
204| `no-outbound-links` | Add at least one link to related docs so a multi-step task doesn't stall on this page. |
205 
206Categories shared with AI View (`long-paragraph`, `context-less-code`, `context-less-table`, `weak-alt-text`, `positional-reference`, `opaque-jsx`, `weak-heading-context`) use the same fixes described on the [AI View](/features/ai-view#fixing-common-issues) page.
207 
208![Run RAG check](/images/run-rag-check-2.png)
209 
210---
211 
212## Run RAG Check across every doc
213 
214You can run a site-wide diagnostic scan across every page in your navigation tree to find low-scoring docs in a single pass:
215 
216```bash
217npx tsx scripts/rag-check-report.ts [threshold]
218```
219 
220The `rag-check-report.ts` script runs the same chunking, question generation, retrieval, and scoring pipeline as the interactive tool against every document in nav.yml. It sorts the results by overall score and outputs pages scoring at or below the specified threshold (default 89).
221 
222For each flagged document, the report displays the overall score, the two lowest-scoring dimensions, and up to five findings:
223 
224```text
225--- /features/example-doc (Example Doc) — overall 72 — weakest: crossLinkQuality=40, retrievability=58
226 [blocking] dead-link: ...
227 [warning] weak-heading-context: ...
228 ...and 3 more findings
229```
230 
231<Callout type="note">
232- Run rag-check-report.ts locally as a diagnostic tool. CamelMind does not wire this script into package.json or CI build workflows.
233- The script skips entries without a source file (such as `noDropdown` group placeholders) or pages that fail to load, counting them separately in the final summary line.
234</Callout>
235 
236---
237 
238## Related tools
239 
240- [AI View](/features/ai-view) — preview the raw text a RAG pipeline ingests, without running chunking or retrieval.
241- [LLMs.txt](/features/llms-txt) — the underlying AI-readable output both tools are built on.

Issues (7)

Warnings (7)