OpenLeafBot
OpenLeafBot is the web crawler operated by Valuesoft ApS for the Open-Leaf service. It collects product manuals and technical documentation that our customers have nominated, so that their own support content can be searched and answered from.
How to identify it
Every request it makes carries this User-Agent header:
OpenLeafBot/1.0 (+https://openleaf.cloud/bot.html)
The product token — the name to use in a robots.txt rule — is
OpenLeafBot. That is the only User-Agent we send: we do
not disguise it as a browser's, and nothing we operate crawls under any other name.
We do render pages in a headless browser engine, so our requests also carry the
ordinary client-hint headers a Chromium-based browser sends — the
User-Agent above is what identifies us.
How to block it
We follow the Robots Exclusion Protocol (RFC 9309).
To stop OpenLeafBot from collecting anything on your site, publish this at
/robots.txt:
User-agent: OpenLeafBot
Disallow: /
To exclude part of a site, name the paths instead:
User-agent: OpenLeafBot
Disallow: /private/
Disallow: /*.pdf$
We read /robots.txt before requesting any page or document, and
re-read it regularly. If the file cannot be reached at all — a server error, a
timeout, or a 429 — we treat the whole site as disallowed until it
can. (A plain 404 is different: it means you have published no rules,
and we proceed.)
Two honest limits. A page we render loads its own images, stylesheets and scripts the way a browser does, and those sub-resources are not checked individually. And if a page we were allowed to request redirects us to one you disallow, the redirect has already been followed by the time we can tell — we discard what came back, never store it, and do not follow any link on it.
We also honour Crawl-delay, up to a ceiling of ten seconds
between requests. If you need us to go slower than that, or to stop
immediately, write to us at the address below and we will act on it
directly rather than waiting for a crawl to notice.
How it behaves
-
A crawl fetches one page at a time and paces itself between requests, slowing
further if your
robots.txtasks it to. Rendering a page also loads that page's own images, stylesheets and scripts, the way a browser does. More than one crawl can be running at a time, so you may occasionally see a few of ours at once. - Each crawl is bounded in advance — a limit on how deep it follows links, how many pages it renders and how many files it collects — so it does not walk a whole site.
- It collects published documents. It does not sign in, and it does not try to reach anything behind a login, a paywall or an access control.
- It follows ordinary links. On a small number of documentation sites whose file lists only appear after using an on-page menu, it operates that menu as a visitor would — so those pages see the same background requests a browser makes.
-
It cannot tell which of your links change something when followed, so it treats
them all alike. Excluding them in
robots.txtis the way to keep us off them — subject to the two limits above.
If you would rather talk to us
If OpenLeafBot has collected something it should not have, is causing load, or you simply want it to stop, email us and say so. We will remove the content and exclude the source. You do not need to justify the request.
Valuesoft ApS · CVR DK26757762
Melanders Vænge 3, 2970 Hørsholm, Denmark
claus@valuesoft.dk