How AISA fetches your pages
The User-Agent AISA sends to your site, what it requests and when, why it does not apply robots.txt Disallow rules, and how to let it through a firewall.
AISA sends requests to your storefront to find your pages and read them. Every one of these requests identifies itself with the same User-Agent:
iAdvize-AISA/1.0 (+https://www.iadvize.ninja/docs/product/site-pages/how-aisa-fetches-pages)The link in it points to this page. The product token, iAdvize-AISA/1.0, does not change without this page changing in the same release, so a firewall rule that matches it keeps working.
What AISA requests, and when
| Request | When |
|---|---|
| Your storefront's address | Once, when you create a site in onboarding or with Create site, to check that it answers and to read the store's name and language. |
/robots.txt | At the start of a sitemap import, when you give no sitemap address. Only its Sitemap: lines are read, from the first 512 KB. |
/sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml | During an import, in that order, only when robots.txt names no sitemap. AISA stops at the first that answers with a sitemap. |
| The sitemaps themselves | During an import: the sitemap found or given, and the sitemaps it lists, on your site only. At most 200 documents per import. |
| A page of your list | When the page is read: as soon as you add it by hand, when you click Read, Read again or Read and generate starters, or once a day by automatic reading if you switched it on. |
/products/{handle}.js (Shopify) | When AISA writes starter questions for a product page and your site has a Shopify catalog connector, to find the product's id in your catalog. |
An import runs when a member starts one from Pages, the REST API or the MCP server, and once during onboarding for a site that never had one. Nothing in AISA watches your pages or re-reads them on a timer, except automatic reading, which is off until you switch it on, and reads each page once.
Page-view counting makes no request to your site. The tag already on your pages reports the page views, and a page found that way is stored without being read.
Reads are bounded per site: 1,000 page reads a day (UTC), whatever started them. See Limits.
How the requests behave
GETrequests overhttps, with no cookie and no login. A sitemap address given ashttpis upgraded tohttps.- AISA reads the HTML your server sends. It does not run JavaScript.
- Redirects are followed one at a time, and each one must stay on your site (its URL or one of its allowed origins). At most 5 per request.
- Time limits: 10 seconds for a page, 15 seconds for a sitemap document or
robots.txt. - A page is read up to 6 MB, a sitemap document up to 50 MB once decompressed. Compressed (
.gz) sitemaps are accepted. - An address that resolves to a private or local network address is never requested.
What is not covered here
The assistant's own web search and web page reading tools go through a third-party search service, not through this fetcher, and do not send this User-Agent.
robots.txt Disallow rules are not applied
AISA reads robots.txt for its Sitemap: lines and nothing else. It does not apply Disallow rules, for any request on this page.
robots.txt tells crawlers that explore a site on their own which paths to skip. AISA does not explore: every page it requests is one you published, asked for, or opted in to.
- Published by you. A sitemap import reads only the sitemaps your site lists or the address you give, and stores the addresses in them without reading them.
- Asked for by you. A page is read when you add it by hand or click Read, or when someone in your organization asks through the API or the MCP server.
- Opted in by you. A page found from shopper traffic is one shoppers viewed on your own storefront. It is read only if you ask, or if you switched on automatic reading for the site.
To keep pages out, use Exclusions rather than robots.txt: a rule such as checkout/* keeps those paths out of every import and every traffic run. Removing a page found by an import or from traffic also keeps it out.
Allow AISA through a firewall
A firewall or bot-protection service (a WAF, a CDN bot rule, a "verify you are human" check) can block AISA. A blocked page ends Failed: with No readable text on the page when a bot check answers instead of the page, or with The read failed. Read again. when the firewall answers with an error. A blocked sitemap import ends with No sitemap found or Sitemap import did not finish.
To let AISA through, add an allow rule that matches the User-Agent:
- Match on the product token: the User-Agent contains
iAdvize-AISA/. Matching the whole string also works, but breaks if the version changes. - Scope the rule to
GETrequests, and if your tool allows it, to the paths AISA needs: your pages,/robots.txtand your sitemap files. - If your tool has a "verified bots" list, AISA is not on it: add a custom rule.
This page lists no IP address range to allow. Match on the User-Agent.
A User-Agent can be copied
Anyone can send this User-Agent. An allow rule based on it lets through any client that copies it, on the paths the rule covers. Keep the rule as narrow as your tool allows, and keep your rate limits on.
After you change the rule, open Pages, select the pages that failed and click Read again, or start a new import.
Pages
List the pages of your site from its sitemap, from shopper traffic or by hand, choose which ones AISA reads, and ask for starter questions written for each page. Questions always arrive switched off, for your review.
Embed the assistant on your storefront
Paste one script tag to load the assistant. Open it from a floating chat bubble, your own button, or both. No login, works from the site key.