This is what an AI/RAG pipeline sees when it indexes this page — the same output served at https://camelmind-docs.vercel.app/api/llms/features/agent-crawler-discoverability.Back to doc

Rendered doc

Agent and Crawler Discoverability

Configure how CamelMind exposes public documentation to AI agents and search crawlers through llms.txt, robots.txt, and sitemap.xml.

CamelMind provides three machine-readable endpoints that help AI agents and search crawlers discover and read your public documentation without scraping rendered HTML:

EndpointPurpose
/llms.txtIndex of public documentation pages, with links to each page's Markdown source.
/robots.txtCrawl rules, AI-bot allowances, and Content Signals policy.
/sitemap.xmlList of public documentation URLs available for search indexing.

CamelMind enables all three endpoints by default and requires no configuration.

This guide explains how CamelMind determines which documentation is public, how it generates each endpoint, and how to configure crawler and AI discoverability.

Note

Each documentation version has its own /llms.txt index at /{version}/llms.txt (for example, /v2/llms.txt). The index uses that version's nav.yml, so agents can discover documentation for a specific version.


Determine public documentation from nav.yml roles

A page is public when its roles field in nav.yml is empty:

yaml
roles: []

CamelMind treats a page with one or more roles as gated, and excludes it from machine-readable discovery endpoints.

lib/nav.ts's getPublicSlugsFromConfig(nav) function collects the public page slugs:

typescript
getPublicSlugsFromConfig(nav) // roles.length === 0

The function walks the complete navigation tree, including:

  • top-level entries
  • nested sections and children
  • noDropdown groups and their direct-link slug
Note

CamelMind does not provide a corresponding getGatedSlugsFromConfig function. sitemap.xml and robots.txt intentionally do not list gated paths.


Exclude documentation from discovery on private sites

Site-wide authentication can override individual page roles. When you enable both auth.enabled: true and auth.requireLogin: true, no documentation page is anonymously reachable, regardless of its roles value.

CamelMind centralizes this rule in lib/agent-discovery.ts:

typescript
getPublicSlugsAcrossVersions(includeVersions)

If the site requires login for every page, this function returns no public documentation slugs.

As a result, sitemap.xml and robots.txt expose only the routes that remain intentionally public on a fully private site.

Note

auth.publicPaths is separate from crawler discoverability. It defines authentication exceptions such as /login and /api/auth; it is not an allowlist of pages that crawlers should discover.


Generate robots.txt with crawl rules and Content Signals

CamelMind implements robots.txt as the app/robots.txt/route.ts Route Handler instead of using Next.js's app/robots.ts metadata convention.

The Route Handler allows CamelMind to include Cloudflare's Content Signals policy:

text
Content-Signal: search=yes, ai-input=yes, ai-train=yes

The standard Next.js app/robots.ts convention supports Allow, Disallow, and Sitemap, but does not provide a field for this Content Signals header.


Generate sitemap.xml from public documentation

app/sitemap.ts uses Next.js's standard MetadataRoute.Sitemap convention — sitemap.xml needs nothing exotic, unlike robots.txt.

The sitemap includes public documentation URLs determined from the navigation configuration and version settings. Gated documentation pages are not included.

API Reference pages generated from OpenAPI specifications are not included because they exist outside nav.yml. Supporting them needs a separate OpenAPI-walking pass.


Keep discovery endpoints accessible on private sites

CamelMind's authentication matcher in proxy.ts excludes robots.txt and sitemap.xml from authentication. These endpoints remain publicly reachable even when the documentation site requires login.

The endpoints determine what information is safe to expose based on the site's authentication and navigation configuration.

Ordinary documentation pages continue to follow the site's authentication rules and redirect to /login when required.


Why gated pages are not listed in robots.txt

CamelMind does not list role-gated documentation paths in robots.txt because robots.txt is publicly accessible.

Listing a gated path would reveal its URL to anonymous visitors without providing meaningful protection. A compliant crawler should already be unable to read the page because page-level RBAC blocks access, while a non-compliant crawler can ignore Disallow rules.

CamelMind applies the same principle to sitemap.xml: omitting gated pages.

If a page must not be accessible to anonymous users, gate the page in nav.yml or require authentication for the entire site. robots.txt and sitemap.xml describe publicly accessible content; they do not provide access control.


Scope discovery by documentation version

By default, sitemap.xml and robots.txt include only versions marked stable: true in versions.yml. Set sitemap.includeVersions to "all" to include unstable and beta versions too.

The unversioned /llms.txt has separate version-resolution behavior: its documentation links come from loadNav(), which resolves unversioned URLs to the first version in versions.yml with stable: true — the same resolution used for bare, unprefixed documentation URLs.

If multiple versions have stable: true, only the first one supplies the unversioned /llms.txt. Each version still has its own versioned /{version}/llms.txt, and the unversioned index links to the available versioned indexes under a "Versions" heading.


Configure AI and crawler discoverability

All three endpoints are configured under the ai key in camelmind.config.ts:

typescript
ai: {
  llmsTxt: {
    enabled: true,
    directive: "For a complete documentation index, see /llms.txt. ...",
  },
  robotsTxt: {
    enabled: true,
    contentSignal: {
      search: "yes",
      aiInput: "yes",
      aiTrain: "yes",
    },
  },
  sitemap: {
    enabled: true,
    includeVersions: "stable",
  },
}
OptionDefaultEffect
robotsTxt.enabledtrueServes a bare Disallow: / without a sitemap reference when set to false.
robotsTxt.contentSignal.search"yes"Signals whether crawlers can use content to build a search index.
robotsTxt.contentSignal.aiInput"yes"Signals whether crawlers can provide content to an AI model to generate a live answer.
robotsTxt.contentSignal.aiTrain"yes"Signals whether developers can use content to train or fine-tune an AI model.
sitemap.enabledtrueServes an empty <urlset> when set to false.
sitemap.includeVersions"stable"Includes unstable and beta documentation versions when set to "all".
Warning

Content Signals are declarations, not access controls. Setting contentSignal.aiTrain to "no" tells cooperating crawlers not to use your content for training, but it cannot technically prevent a non-compliant scraper or an agent that fetches the page directly. Use Content Signals to communicate your preference; apply authentication and RBAC to protect confidential content.


Choose access control or Content Signals

Use authentication and RBAC when you need to prevent access. Use Content Signals when you want public content to remain discoverable while communicating how cooperating crawlers may use it.

GoalConfiguration
Prevent automated and anonymous access to confidential or restricted contentAdd a roles entry in nav.yml, or require authentication for the entire site.
Keep documentation publicly readable but discourage AI trainingLeave the page public and set contentSignal.aiTrain to "no".

If it would be unacceptable for a non-compliant agent to access or train on the content, do not rely on Content Signals alone. Gate the content with authentication and RBAC.


Verify discovery endpoints locally

bash
npm run dev
curl http://localhost:3000/robots.txt
curl http://localhost:3000/sitemap.xml
curl http://localhost:3000/llms.txt

To check behavior on a fully private site:

bash
CAMELMIND_AUTH_ENABLED=true CAMELMIND_AUTH_REQUIRE_LOGIN=true npm run dev

On a fully private site, sitemap.xml should contain only the routes that remain intentionally public, such as / and /home. robots.txt should disallow the private documentation while preserving access to those public routes and /llms.txt.


  • llms.txt — Learn how CamelMind generates AI-readable documentation indexes.
  • Authentication & RBAC — Configure the roles field and site-wide authentication that determine documentation access and discoverability.

What the AI sees

1> For a complete documentation index, see /llms.txt. To read any public page as Markdown, append .md to the URL.
2 
3CamelMind provides three machine-readable endpoints that help AI agents and search crawlers discover and read your public documentation without scraping rendered HTML:
4 
5| Endpoint | Purpose |
6| --- | --- |
7| `/llms.txt` | Index of public documentation pages, with links to each page's Markdown source. |
8| `/robots.txt` | Crawl rules, AI-bot allowances, and Content Signals policy. |
9| `/sitemap.xml` | List of public documentation URLs available for search indexing. |
10 
11CamelMind enables all three endpoints by default and requires no configuration.
12 
13This guide explains how CamelMind determines which documentation is public, how it generates each endpoint, and how to configure crawler and AI discoverability.
14 
15<Callout type="note">
16Each documentation version has its own `/llms.txt` index at `/{version}/llms.txt` (for example, `/v2/llms.txt`). The index uses that version's `nav.yml`, so agents can discover documentation for a specific version.
17</Callout>
18 
19---
20 
21## Determine public documentation from nav.yml roles
22 
23A page is public when its `roles` field in `nav.yml` is empty:
24 
25```yaml
26roles: []
27```
28 
29CamelMind treats a page with one or more roles as gated, and excludes it from machine-readable discovery endpoints.
30 
31`lib/nav.ts`'s `getPublicSlugsFromConfig(nav)` function collects the public page slugs:
32 
33```typescript
34getPublicSlugsFromConfig(nav) // roles.length === 0
35```
36 
37The function walks the complete navigation tree, including:
38 
39* top-level entries
40* nested sections and children
41* `noDropdown` groups and their direct-link `slug`
42 
43<Callout type="note">
44CamelMind does not provide a corresponding `getGatedSlugsFromConfig` function. `sitemap.xml` and `robots.txt` intentionally do not list gated paths.
45</Callout>
46 
47---
48 
49## Exclude documentation from discovery on private sites
50 
51Site-wide authentication can override individual page roles. When you enable both `auth.enabled: true` and `auth.requireLogin: true`, no documentation page is anonymously reachable, regardless of its `roles` value.
52 
53CamelMind centralizes this rule in `lib/agent-discovery.ts`:
54 
55```typescript
56getPublicSlugsAcrossVersions(includeVersions)
57```
58 
59If the site requires login for every page, this function returns no public documentation slugs.
60 
61As a result, `sitemap.xml` and `robots.txt` expose only the routes that remain intentionally public on a fully private site.
62 
63<Callout type="note">
64`auth.publicPaths` is separate from crawler discoverability. It defines authentication exceptions such as `/login` and `/api/auth`; it is not an allowlist of pages that crawlers should discover.
65</Callout>
66 
67---
68 
69## Generate robots.txt with crawl rules and Content Signals
70 
71CamelMind implements `robots.txt` as the `app/robots.txt/route.ts` Route Handler instead of using Next.js's `app/robots.ts` metadata convention.
72 
73The Route Handler allows CamelMind to include Cloudflare's Content Signals policy:
74 
75```text
76Content-Signal: search=yes, ai-input=yes, ai-train=yes
77```
78 
79The standard Next.js `app/robots.ts` convention supports `Allow`, `Disallow`, and `Sitemap`, but does not provide a field for this Content Signals header.
80 
81---
82 
83## Generate sitemap.xml from public documentation
84 
85`app/sitemap.ts` uses Next.js's standard `MetadataRoute.Sitemap` convention — sitemap.xml needs nothing exotic, unlike `robots.txt`.
86 
87The sitemap includes public documentation URLs determined from the navigation configuration and version settings. Gated documentation pages are not included.
88 
89API Reference pages generated from OpenAPI specifications are not included because they exist outside `nav.yml`. Supporting them needs a separate OpenAPI-walking pass.
90 
91---
92 
93## Keep discovery endpoints accessible on private sites
94 
95CamelMind's authentication matcher in `proxy.ts` excludes `robots.txt` and `sitemap.xml` from authentication. These endpoints remain publicly reachable even when the documentation site requires login.
96 
97The endpoints determine what information is safe to expose based on the site's authentication and navigation configuration.
98 
99Ordinary documentation pages continue to follow the site's authentication rules and redirect to `/login` when required.
100 
101---
102 
103## Why gated pages are not listed in robots.txt
104 
105CamelMind does not list role-gated documentation paths in `robots.txt` because `robots.txt` is publicly accessible.
106 
107Listing a gated path would reveal its URL to anonymous visitors without providing meaningful protection. A compliant crawler should already be unable to read the page because page-level RBAC blocks access, while a non-compliant crawler can ignore `Disallow` rules.
108 
109CamelMind applies the same principle to `sitemap.xml`: omitting gated pages.
110 
111If a page must not be accessible to anonymous users, gate the page in `nav.yml` or require authentication for the entire site. `robots.txt` and `sitemap.xml` describe publicly accessible content; they do not provide access control.
112 
113---
114 
115## Scope discovery by documentation version
116 
117By default, `sitemap.xml` and `robots.txt` include only versions marked `stable: true` in `versions.yml`. Set `sitemap.includeVersions` to `"all"` to include unstable and beta versions too.
118 
119The unversioned `/llms.txt` has separate version-resolution behavior: its documentation links come from `loadNav()`, which resolves unversioned URLs to the first version in `versions.yml` with `stable: true` — the same resolution used for bare, unprefixed documentation URLs.
120 
121If multiple versions have `stable: true`, only the first one supplies the unversioned `/llms.txt`. Each version still has its own versioned `/{version}/llms.txt`, and the unversioned index links to the available versioned indexes under a "Versions" heading.
122 
123---
124 
125## Configure AI and crawler discoverability
126 
127All three endpoints are configured under the `ai` key in `camelmind.config.ts`:
128 
129```typescript
130ai: {
131 llmsTxt: {
132 enabled: true,
133 directive: "For a complete documentation index, see /llms.txt. ...",
134 },
135 robotsTxt: {
136 enabled: true,
137 contentSignal: {
138 search: "yes",
139 aiInput: "yes",
140 aiTrain: "yes",
141 },
142 },
143 sitemap: {
144 enabled: true,
145 includeVersions: "stable",
146 },
147}
148```
149 
150| Option | Default | Effect |
151| --- | --- | --- |
152| `robotsTxt.enabled` | `true` | Serves a bare `Disallow: /` without a sitemap reference when set to `false`. |
153| `robotsTxt.contentSignal.search` | `"yes"` | Signals whether crawlers can use content to build a search index. |
154| `robotsTxt.contentSignal.aiInput` | `"yes"` | Signals whether crawlers can provide content to an AI model to generate a live answer. |
155| `robotsTxt.contentSignal.aiTrain` | `"yes"` | Signals whether developers can use content to train or fine-tune an AI model. |
156| `sitemap.enabled` | `true` | Serves an empty `<urlset>` when set to `false`. |
157| `sitemap.includeVersions` | `"stable"` | Includes unstable and beta documentation versions when set to `"all"`. |
158 
159<Callout type="warning">
160Content Signals are declarations, not access controls. Setting `contentSignal.aiTrain` to `"no"` tells cooperating crawlers not to use your content for training, but it cannot technically prevent a non-compliant scraper or an agent that fetches the page directly. Use Content Signals to communicate your preference; apply authentication and RBAC to protect confidential content.
161</Callout>
162 
163---
164 
165## Choose access control or Content Signals
166 
167Use authentication and RBAC when you need to prevent access. Use Content Signals when you want public content to remain discoverable while communicating how cooperating crawlers may use it.
168 
169| Goal | Configuration |
170| ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |
171| Prevent automated and anonymous access to confidential or restricted content | Add a `roles` entry in `nav.yml`, or require authentication for the entire site. |
172| Keep documentation publicly readable but discourage AI training | Leave the page public and set `contentSignal.aiTrain` to `"no"`. |
173 
174If it would be unacceptable for a non-compliant agent to access or train on the content, do not rely on Content Signals alone. Gate the content with authentication and RBAC.
175 
176---
177 
178## Verify discovery endpoints locally
179 
180```bash
181npm run dev
182curl http://localhost:3000/robots.txt
183curl http://localhost:3000/sitemap.xml
184curl http://localhost:3000/llms.txt
185```
186 
187To check behavior on a fully private site:
188 
189```bash
190CAMELMIND_AUTH_ENABLED=true CAMELMIND_AUTH_REQUIRE_LOGIN=true npm run dev
191```
192 
193On a fully private site, `sitemap.xml` should contain only the routes that remain intentionally public, such as `/` and `/home`. `robots.txt` should disallow the private documentation while preserving access to those public routes and `/llms.txt`.
194 
195---
196 
197## Related documentation
198 
199* [llms.txt](/features/llms-txt) — Learn how CamelMind generates AI-readable documentation indexes.
200* [Authentication & RBAC](/features/auth-rbac) — Configure the `roles` field and site-wide authentication that determine documentation access and discoverability.

Issues (1)

Warnings (1)