iAdvizeDocs
Pages

Pages

List the pages of your site from its sitemap, from shopper traffic or by hand, choose which ones AISA reads, and ask for starter questions written for each page. Questions always arrive switched off, for your review.

Pages is where you tell AISA which pages of your site to learn from. Pages reach the list three ways: a sitemap import, shopper traffic, and addresses you add by hand. AISA reads a page, then writes starter questions for the pages you pick. Every generated question lands switched off in the library. Nothing reaches a shopper until you read it and switch it on.

Open Knowledge → Pages in the dashboard sidebar for a site, or go to /dashboard/sites/{siteId}/pages.

What a page is

A page is one address on your site that AISA has stored. Only two kinds of page can produce questions:

  • Product pages.
  • Category pages.

The read decides the kind. A Home page, a Content page (a size guide, a policy) or a Blog page (an article or a blog post) is kept in the list as Unsupported and produces nothing: a blog page is treated as content. A page AISA cannot classify is treated like a content page too.

A page found by a sitemap import or from traffic gets a type guessed from its address as soon as it is found, shown with guessed under it until the page is read. The first segment of the path decides, after any language segment such as /fr/:

Path starts withGuessed type
Nothing (/, or a language segment alone)Home
/products/, /product/, /produits/, /produit/ or /p/, then moreProduct
One of those alone (/products), or /collections/, /collection/, /category/, /categories/, /c/Category
/blogs/, /blog/, /news/ or /articles/Blog
A content segment such as /pages/, /policies/, /legal/, /about/, /contact/, /faq/ or /help/Content
Anything elseNo guess

A guess only describes the page: it decides nothing about reading, and the read replaces it. A read that reaches the page decides its type, even when it ends Unsupported; a read that fails keeps the guess. A page whose address matches no row, and a page you add by hand (it is read at once), get no guess: the type shows Not known yet until the page is read.

When it reads a page, AISA recognizes a blog post from a Shopify /blogs/<blog>/<article> address, Article, BlogPosting or NewsArticle structured data, or the body class WordPress gives a single post. An og:type of article counts only on an address that already points at a blog (/blog/..., /news/...): some SEO plugins set it on every page, static pages included.

Where pages come from

SourceHow a page gets thereRead when
SitemapA sitemap import found it.You ask, or automatic reading
Shopper trafficShoppers viewed it at least 3 times over 30 days. See Automation.You ask, or automatic reading
Added by handYou added its address.At once

A page keeps the source that added it first. A page found by a sitemap import or from traffic waits as Not read: AISA has stored its address and nothing else. It is read only when you ask, or when automatic reading picks it.

Import from the sitemap

Click Import from sitemap (on an empty site, the empty state offers the same). AISA finds your sitemap and lists your pages as Not read. It reads none of them.

Your first import may already have run

During onboarding, the site analysis starts a sitemap import on its own, for a site that never had one. If your Pages list is already full when you first open it, that is why.

Where AISA looks for the sitemap

Leave Sitemap address empty and AISA looks in this order:

  1. The Sitemap: lines of your robots.txt. Every line that points to your site is used. A line that points to another host is dropped.
  2. If robots.txt names none, the first of these that answers with a sitemap: /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml.

Or give the address yourself. It must be on your site's URL or on one of its allowed origins. An address elsewhere is refused before any request is made to it: "This address is not on your site or an origin it allows." An http address is upgraded to https.

AISA reads robots.txt for its Sitemap: lines only. It does not apply its Disallow rules. See How AISA fetches your pages for why, and for how to let AISA through a firewall.

What an import does

  • A sitemap index is followed to the sitemaps it lists, two levels deep at most (index, index, then the list of pages). A child sitemap on another host is not requested. Compressed (.gz) sitemaps are read.
  • Each address is cleaned like an address you type (see Add pages by hand) and must be on your site. A link to a file that is not a web page (a PDF, an image, a .xml, .json or .md file, a video, a font) is left out.
  • An address already in your list is kept as it is. An excluded address is skipped.
  • A new page is stored as Not read, with a type guessed from its address and a language guessed from the sitemap's hreflang for that address, or else from a language segment at the start of its path (/fr/...). The list marks each guess as such.
  • The import runs in the background. You can leave the page. Pages appear as AISA finds them.

One import runs at a time per site. Starting a second one while the first runs is refused: "An import is already running. Wait for it to finish." An import that makes no progress for 30 minutes is shown as failed and stops blocking a new one.

An import stops at these bounds:

BoundValue
Sitemap documents per import200
Addresses taken per import50,000
Size of one sitemap document50 MB once decompressed (the sitemap protocol's own maximum)
Time per request15 seconds
Pages a site holds50,000, all sources together (see Limits)

The last-import summary

The description under the Pages title ends with the state of the last import: Importing your sitemap, n pages so far with a spinner, Sitemap imported with its date, Sitemap import did not finish with its date, or No sitemap found. For a member who can edit, the first two finished states are followed by Import again and the last by Give its address; both open the import dialog. Before a site's first import, the description says only what AISA does with your pages and when it reads them.

The Import from sitemap dialog sums up the last import: how many pages it added, how many were already in your list, how many addresses were left out and why (not on your site, not a page, not valid, containing *, too long, over the page limit), and how many it skipped because you excluded them.

Missing from the sitemap

When an import completes cleanly (every sitemap document read, no bound reached), every page from the sitemap that it did not meet is marked Missing from the sitemap since the date of that import. The mark is cleared when a later import finds the page again. Nothing is removed: the mark is there so you can decide. Filter on it with Missing from sitemap, then remove the pages you no longer want.

Pages found from traffic or added by hand are never marked.

Pages found from traffic

When page-view counting is on (it is by default), a page of your site viewed at least 3 times over the last 30 days that the list does not hold yet is added once a day, as Not read, with the source Shopper traffic. This catches pages your sitemap misses. It is never read on its own. See Automation.

Add pages by hand

Click Add pages (on an empty site, the empty state offers the same). Paste one address or path per line, then click Add n pages.

  • A path such as /products/linen-shirt is resolved against your site's URL (a site with no URL accepts full addresses only). A full address must be http or https and sit on your site's own origin: the one in your site's URL or one of its allowed origins. An http address is upgraded to https. When an address is off your site but its host differs from an allowed origin only by a leading www. (added or removed), it is rewritten to that origin. This only works toward an origin listed exactly, never toward a wildcard origin.
  • The dialog judges every line before you submit and names the outcome of each: New, New, cleaned to {path}, Already added, ignored, Same as line n, ignored, or a refusal (not on your site, not a valid address, contains *, a path and query longer than 200 characters, or "The site holds its limit of 50,000 pages"). A refused line never blocks the others.
  • Addresses are cleaned: the fragment (#...) is removed, and so are tracking parameters: every utm_ parameter, plus gclid, gbraid, wbraid, dclid, fbclid, msclkid, ttclid, twclid, li_fat_id, igshid, yclid, srsltid, mc_cid, mc_eid, _ga, _gl and _kx. A trailing slash is removed, except on the home page (/). Any other query parameter (ref or variant, for example) is kept as written, in its order, and the path keeps its case. An http upgrade or a www. rewrite also counts as cleaning.
  • One request holds at most 50 non-empty lines. Above that, nothing is added and the dialog says how many lines you sent.
  • Adding an excluded address by hand creates the page as usual and lifts that address's exclusion.

AISA starts reading pages you add by hand as soon as you add them. A page that does not answer is still added: it shows as Failed, with the reason, so you can read it again later.

Read pages

A Not read page has two actions, in its row's ⋯ menu, on the page's own view, and in the bulk bar when every selected page is Not read:

  • Read: AISA reads the page. Nothing is generated.
  • Read and generate starters: AISA reads each page first, then writes questions for each one that ends Read, in one request. The dialog lists the pages, shows how many reads and generations are left today, and offers the same Replace option as Generate starters. When the selection does not fit in what is left of today's reading limit, it says how many will be read: "Only n of m pages fit in today's reading limit. The others are not read." A page that fails, or is not a product or category page, gets no questions. This action needs Starter questions: Edit too, and is not shown without it.

A page that is already Read, Failed or Unsupported has Read again instead. Reading again replaces the text, type and language AISA held for it.

To have pages read without asking, switch on automatic reading under Site Settings → Automation. It is off by default. Once a day, it reads the Not read pages that reach your threshold, most viewed first. With a threshold of 0, every Not read page is eligible, still most viewed first, or oldest found first when the site does not count page views.

Every read counts against the site's daily reading limit. A page you ask for after the limit is spent ends Failed with "Daily reading limit reached". Read it again after the reset. A page automatic reading could not read (the limit was spent, or the run could not start) goes back to Not read instead, and is considered again the next day.

Statuses

The line under the header counts your pages, how many are read and how many failed. Near the page limit it also shows how many pages the site holds out of 50,000. The Status column and filter use five:

StatusWhat it means
Not readFound by a sitemap import or from traffic, and not read yet. Read on request or by automatic reading.
ReadingWaiting to be read or being read now. A page you add by hand starts here.
ReadAISA stored the page's text. A product or category page can now be generated for.
FailedThe read did not work. The row names the reason, and Read again retries it.
UnsupportedThe page was read but cannot produce questions. The row names the reason.

A page that stays in Reading for more than 30 minutes is shown as Failed ("The read failed. Read again."), so a lost job never leaves a page reading forever.

Reasons you can see on a row:

ReasonStatus
This address is not allowedFailed
The page took too long to answerFailed
Page not found (404)Failed
The page redirects to another siteFailed
No readable text on the pageFailed
The read could not start. Read again.Failed
The read failed. Read again.Failed
Daily reading limit reachedFailed
Home pages are not supported yetUnsupported
Content pages are not supported yetUnsupported
This site does not serve the page's languageUnsupported
Duplicate of {path}Unsupported

No readable text on the page is what a page that builds its content with JavaScript, or one that answers a bot check, usually returns: AISA reads the HTML the server sends and does not run scripts. A page it reads must also stay on your site: a redirect to another host fails with The page redirects to another site. A redirect to an address that contains * or whose path and query exceed 200 characters fails with This address is not allowed, because no question could target it.

Duplicates and the canonical address

When AISA reads a page, it stores the canonical address the page declares (<link rel="canonical">), cleaned like any address and only when it is on your site. A canonical on another domain is ignored and the page is judged on its own.

When the canonical is the address of another page in your list, the page ends Unsupported with Duplicate of {path}, naming that page. Both pages stay in the list (nothing is merged), and the duplicate is never generated for. A variant address such as /products/oak-table?variant=2 whose canonical is /products/oak-table is the usual case. When the canonical page is not in your list yet, the page is judged on its own; reading it again after the canonical page is added resolves it.

The page view shows the Canonical address and, for a duplicate, the page it duplicates.

The list

The table shows 50 pages at a time, sorted by most viewed over the last 30 days by default. When page-view counting is off for the site, the list has no views: it is sorted newest first by default, and the Views (30 d) column is not shown. Each row has:

  • Page: the path and title, or "Not read yet" for a page not read yet, and Missing from the sitemap since a date when it applies.
  • Views (30 d): the page's views over the last 30 days, from page-view counting, recomputed once a day. Shown only while the site counts page views.
  • Source: Sitemap, Shopper traffic or Added by hand.
  • Type and Language, each marked guessed until the page is read: the type is guessed from the address (see What a page is), the language from the sitemap or the address.
  • Status, Starters (how many questions it produced and how many are live, or Generating starter questions while a generation runs for the page) and Last read.

Sort the list

Click the Page, Views (30 d) or Last read header to sort by that column. Click the same header again to reverse it. The sorted column shows an arrow pointing up (ascending) or down (descending). The other columns do not sort.

  • Page sorts by address: A to Z first, then Z to A.
  • Views (30 d) sorts by views: most viewed first, then least viewed first. It is the default sort, so a click on it while it is sorted shows the least viewed first. It is offered only while the site counts page views.
  • Last read sorts by the date of the last read: most recent first, then oldest first. Pages never read stay at the end in both directions.

Changing the sort, a filter or the search goes back to the first page of results. The sort lives in the page's address, so a copied link reopens the same order.

Filter and act

  • Search matches the address or the title.
  • Filters: status, type (a page not read yet matches on its guessed type), language (the language read, or the guess for a page not read yet), source, Missing from sitemap, and starters (With starters or No starters yet).
  • Select pages with the checkboxes to act on several at once. You can select 50 at most.
  • Bulk actions: Generate starters (or Read and generate starters when every selected page is Not read), Read again (or Read), and Remove. Each row also has these in its ⋯ menu, plus Open.

Click a row to open the page.

The list updates itself

While something is in progress on the site (a page waiting to be read or being read, a starter generation, or a sitemap import), the list refreshes itself 3 seconds after its last refresh finished: statuses change, new pages appear and starter counts fill in without a reload. It stops as soon as nothing is in progress, and pauses while the browser tab is hidden. A refresh keeps an open dialog, your selection and your scroll position.

A page shows Generating starter questions, with a spinner, from the moment you ask for questions until the generation is recorded in its history. After 30 minutes with no result, it stops showing as generating.

Exclusions

Exclusions (the control next to Import from sitemap, with a count) lists what a sitemap import and shopper traffic must never add back. They apply to discovery only: a page you add by hand is never excluded and never removed by a rule.

There are two kinds:

  • Removed addresses. Removing a page that came from the sitemap or from traffic records its address here, so it does not come back. The remove dialog says so: "Pages found in your sitemap or from traffic will not come back from either." Allow again lifts one, and so does adding that address by hand. Removing a page you added by hand records nothing. The sheet shows the 100 most recent.
  • Rules. A path pattern, written without the first slash, where * matches any text, the same syntax as starter question URL targeting. blogs/* keeps out everything under /blogs/. */cart keeps out any path ending in /cart. A rule matches the path only, never the query string, and is case-sensitive. A leading / is accepted and removed. A rule can't contain a space, ? or #, and is at most 200 characters.

Before you save a rule, the dialog says how many pages it will remove. Saving it removes, at once, every page from the sitemap or from traffic whose path matches. Pages you added by hand stay. The starter questions of removed pages stay in Questions. A site holds at most 50 rules.

Removing a rule or allowing an address again restores nothing at once: the next import or the next day's traffic run can find those pages again.

Every member who can open Pages can see the exclusions. Adding and removing them needs Site pages: Edit.

The page view

/dashboard/sites/{siteId}/pages/{pageId} shows one page.

  • Starter questions from this page: while a generation runs for the page, Generating starter questions with a spinner. Then each generated question with its type, whether it is live, and its impressions, clicks and click rate over the last 14 days, the same figures and verdicts as in Questions. Review in Questions opens the library.
  • What AISA read: the text AISA stored for the page, shown as plain text. It feeds the question writer. The page view says whether answers were checked against your catalog record or against this text (see Evidence). A Not read page says where it was found and that AISA reads it when you ask, or by automatic reading.
  • History: every read and every generation, newest first, ending with how the page was added ("Added by hand.", "Found in the sitemap import." or "Found from shopper traffic."). A generation says how many questions it wrote and how many candidates did not pass the check.
  • Details: the address, which opens the page itself in a new tab, the path the questions are shown on, its canonical address, the page it duplicates, its type and language (each marked "guessed from the address" before a read), how and when it was added, its views over 30 days and when they were counted, whether it is in the sitemap (last seen, last modified, missing since), which catalog record was used, and when it was last read.

While the page is being read or its questions written, the page view refreshes itself the same way the list does, so Read again and Generate starters show their result without a reload.

Generate starter questions

Select one or more Read pages and click Generate starters (or use a row's ⋯ menu, or the link in the Starters column). The dialog lists each page with its type and language and shows how many generations are left today. For pages not read yet, use Read and generate starters.

  • AISA writes up to 5 questions per page, in the language of the page, and shows each one on that page only: the question targets the page's own path and query.
  • Only Read product and category pages qualify. A page that does not is listed with its reason and skipped.
  • Generation runs in the background. You can leave the page. New questions appear in Questions as soon as they are written.
  • Every question is created switched off. You review it, edit it if you like, and switch it on.
  • Replace earlier generated questions that are still switched off is ticked by default. When you generate again for a page, the questions generated earlier for that page that are still off are deleted when new ones are written. A question you switched on, a question you edited, and a question you wrote yourself are never touched. Untick it to keep the old ones beside the new ones.

A page can end a generation with no question: No starter passed the check appears in its history. That is a normal outcome. Each candidate goes through automatic checks, including whether it can be answered from the evidence below, and a page whose candidates all fail produces none.

Languages

AISA writes in the page's language when your site serves it: English, French, German, Italian, Spanish or Dutch. A page in a language your site has not enabled is stored as Unsupported ("This site does not serve the page's language") and is never generated for. A page that declares no language is written in your site's default language.

Evidence: what the questions are checked against

A starter question is only worth showing if the assistant can answer it. So each candidate is checked against what the assistant would know.

  • Catalog record. When your site has a catalog connector AISA can query (Shopify, Adobe Commerce, and some custom connectors) and the page's product is found there, questions are checked against that record: title, price, availability, variants and description.
  • Page text. When there is no usable connector, or the product is not found, the questions are checked against the text AISA read from the page. The page view and the question's own view say so: "Checked against the page text, not your catalog."

A question checked only against the page text may not be answerable yet

A question checked against the page text is checked against what the page says, not against what the assistant can look up. The assistant may not be able to answer it in a conversation until it can read your pages. The question's own view carries that note: "The assistant may not be able to answer it until it can read your pages." Read these questions with that in mind before you switch them on.

For a category page, the catalog check is a topical sample found by the page's title, not the collection's full list.

Review the questions

Generated questions show up in Questions with an AI badge in the new Type column, and a notice above the grid says how many are waiting for review. See Generated questions.

Automation

Under Site Settings → Automation you can let AISA work on its own. See Automation.

  • Count page views (on by default) ranks your pages and finds the ones your sitemap misses.
  • Read pages automatically (off by default) reads, once a day, the Not read pages that reach your threshold, most viewed first.
  • Generate starter questions automatically (off by default) writes questions for every page that is read, by hand or automatically.

Limits

Each site has a daily allowance. It counts in UTC and resets at midnight UTC, not at midnight in your time zone.

Prop

Type

What is kept, and for how long

  • The page text is stored with the page. AISA keeps the text it read (up to 8,000 characters) so you can see what the questions were written from. It is deleted when you remove the page, delete the site, or the organization is deleted.
  • Removing a page keeps its questions. The questions stay in the library. They lose only their link to the page: their banner reads "The page it was written from was removed."
  • Removing a page and adding it again duplicates its questions. The re-added page is a new page with no link to the old questions. Generating again for it does not replace them, since replacement only touches questions linked to the page. Delete the old ones in Questions if you do not want both.
  • Questions are not regenerated when a page changes. Nothing watches your pages. Read again refreshes the stored text only. If automatic generation is on, a page that finishes reading gets questions again, replacing the untouched ones that are still off. Otherwise, generate again yourself.
  • Page-view counts are kept 90 days. See Counting page views.
  • Pages are not shown to shoppers. Only the questions you switch on are.

Pages are read by AISA, not by the assistant's Web page reading tool. The domain policy on the Tools page does not apply to them. See How AISA fetches your pages for what is requested from your site, and when.

Permissions

Viewing the list, the page view and the exclusions requires Site pages: View. Importing, adding, reading, reading again, removing pages and managing exclusions requires Site pages: Edit. Generating, and Read and generate starters, also require Starter questions: Edit, because they write to the question library. Without Site pages: View (or Edit), Pages does not appear in the sidebar and opening the URL shows a restricted-access message. See Roles.

Over the API and MCP

Everything on this page is also available over the REST API and the MCP server, under the same rules: the same limits, the same eligibility, and questions that always arrive switched off. The API's page list is sorted newest first unless you ask for another sort, and every sort runs in both directions with order.

Audit log

Starting an import, adding, reading, reading again, removing, asking for questions, adding an exclusion rule and removing exclusions each write one audit entry with counts only, never an address, a rule or a text. Automatic reading and automatic generation write one entry per run, attributed to System. Pages found by an import or from traffic get no entry of their own: the import request is the recorded decision, and the traffic run is a system job. The import that onboarding starts writes no entry.

On this page