Open to machines
AI and crawler policy
AI crawlers and assistants are welcome here. Every public page on this site may be fetched, indexed, summarised, quoted and used to answer someone's question. There is no separate agreement to sign and no key to request.
This page exists so that position is readable by a person. The rule itself is in robots.txt, and both are generated from one registry in this repository, so the page you are reading cannot describe a permission the server does not actually grant.
Where to start
- /llms.txt: the site index written for language models: every public page with a sentence describing it, plus the facts worth knowing before quoting a number from here.
- /sitemap.xml: every indexable URL, including one page per country with data.
- /robots.txt: the machine-readable rules this page explains.
- /methodology: what this site refuses to blend, guess or convert. Read it before treating two figures here as comparable.
- /mcp-server. There is also a Model Context Protocol endpoint, but unlike the pages it is authenticated: it reads one organisation's own workspace, so it needs a personal access token issued from that account. It is not an open data API and nothing in it is reachable by crawling.
Agents named explicitly
The wildcard rule already permits all of these. They are named one by one because some operators only obey a group that names them, and because an unnamed agent has no way to tell an intentional welcome from an oversight.
Crawlers and indexers
These fetch in bulk to build an index or a training corpus. Nobody is waiting on the response.
GPTBotOpenAIOAI-SearchBotOpenAIClaudeBotAnthropicClaude-WebAnthropicanthropic-aiAnthropicClaude-SearchBotAnthropicPerplexityBotPerplexityMeta-ExternalAgentMetaCCBotCommon CrawlGoogle-ExtendedGoogleApplebot-ExtendedAppleBytespiderByteDanceAmazonbotAmazoncohere-aiCohere
Assistants fetching for a person
These are an AI acting for someone who asked for this page, right now. Blocking one would not be a decision about training data. It would be refusing to serve a visitor who happens to be reading through an assistant.
ChatGPT-UserOpenAIClaude-UserAnthropicPerplexity-UserPerplexityDuckAssistBotDuckDuckGoMistralAI-UserMistral AIMeta-ExternalFetcherMeta
An agent that is not on this list is still allowed. The wildcard rule permits every user agent; nothing here is an allow-list.
What is closed, and why
These paths are disallowed for every user agent, AI or not: /api/, /app/, /admin/, /sign-in, /sign-up, /legal-acceptance, /no-access.
None of that is a secret and none of it is a security measure. The signed-in workspace and the admin console refuse an unauthenticated request on the server, and the API endpoints authenticate and rate-limit on their own terms. They are listed so a crawler does not spend its budget on pages that will answer it with a redirect to a sign-in form.
A crawler and an account see different things
This matters more here than on a plain content site, because this one has both a public half and a private one.
What a crawler sees is the public half: food-market observations ingested daily from named official providers, each with its source URL, the date it was observed and the time Novus fetched it, plus the pages that explain how those figures are normalised and compared. It is aggregated provider data. It contains no personal information about anyone, so crawling this site collects none. There is none there to collect.
What an account holds is a restaurant's own operating data: uploaded supplier invoices and purchase ledgers, menu and recipe costings, review exports, supplier match decisions and saved reports. That is a tenant's private business data. It lives behind /app, it is scoped to one organisation on the server, it is never rendered on a public page, and no crawler has ever been able to reach it. A figure you find on this site is market data; it is never a customer's numbers.
What this site does collect from ordinary visitors is set out in the privacy notice: server and security logs process IP address, request metadata, user agent and timestamps for delivery and abuse prevention. Google Analytics is not requested at all until analytics consent is explicitly granted, and advertising follows separate marketing consent. A crawler that requests pages and stores no cookie triggers neither.
What we ask in return
This is a request, not a condition, and access does not depend on it. Nothing below is enforced and nothing below changes what robots.txt permits.
- Name the source and link to the page. If an answer uses a figure from here, a link lets the reader check it, and checking is the entire point of a site that publishes provenance beside every number.
- Carry the provenance with the figure. Every observation here has an originating provider, an observation date and a source URL. A number repeated without them is less true than the number was.
- Do not present a figure as current without checking when it was observed. Sources publish on their own cadence, a provider that is failing still returns its last successful observation, and the ingestion schedule says when each one is expected.
- Do not invent the numbers this site refuses to state. Where two comparable observations do not exist, the product shows an empty state rather than a computed metric, and it never converts between currencies because the data carries no exchange rate. An absent number here is a deliberate statement.
- Fetch at a reasonable rate. Public API routes are rate-limited and will say so in their response headers rather than failing silently.
Contact
Questions about AI access, a crawl that is being refused, or a correction to something an assistant reported from this site: support@novusstreamsolutions.com.