Crawlers and search

robots.txt, the sitemap, llms.txt, security.txt and the Link header.

robots.txt

All crawlers, including AI crawlers, may read the site. Two API paths are disallowed because they only accept submissions. A Content-Signal line states how the content may be used:

User-Agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /
Disallow: /api/v1/agents/
Disallow: /api/v1/audit-requests

Content Signals

search=yes, ai-input=yes, ai-train=no: content may be indexed for search and used as input to AI answers, not used to train models. The same value is sent as a Content-Signal header on every page.

sitemap.xml

/sitemap.xml lists the public pages of duskstate.dev. These docs have their own at docs.duskstate.dev/sitemap.xml.

llms.txt

/llms.txt is a short summary for language models: what Dusk State is, the labels, the main pages and the key machine-readable files. It repeats the rule that only capabilities in /capabilities.json are available. /llms-full.txt carries the full text of the docs and policies.

security.txt

/.well-known/security.txt (RFC 9116) names security@duskstate.dev as the contact and links the corrections page as the policy.

Every API response carries a Link header pointing to the OpenAPI description (service-desc), the publisher file (describedby), status (status) and the API policy (service-doc). A client that lands on any endpoint can find the rest.