Assist
Content Ingestion
Assist answers only from content you give it. You publish a JSON index of your content at a URL; Assist polls it hourly and re-indexes whenever it changes. Your build pipeline is the source of truth.
The Index File
A single JSON document, deliberately small — most site generators can emit it in a few dozen lines of build script:
{
"hash": "3f9c2a71b4e8d05c",
"chunks": [
{
"url": "/pricing",
"title": "Pricing",
"section": "pricing",
"text": "Storage is billed at $0.015 per GB per month. There are no egress fees within…"
},
{
"url": "/docs/storage/buckets",
"title": "Buckets",
"section": "docs",
"text": "A bucket is a flat namespace for objects. Names must be globally unique and…"
}
]
} | Field | Notes | |
|---|---|---|
hash | Required | Any string that changes when the content changes (a content hash is ideal). Assist re-fetches the index hourly but only re-indexes when the hash differs — and a changed hash also invalidates the answer cache. |
chunks[].url | Required | The page the text came from, as a relative path (e.g. /pricing). Used for citations and search results. |
chunks[].title | Recommended | The page title, shown in search results and citation chips. |
chunks[].section | Optional | A coarse grouping label (e.g. docs, pricing, blog). Shown in search results. |
chunks[].text | Required | The chunk content: plain text, no HTML. |
Chunking Guidance
A chunk is the unit of retrieval: when a visitor asks a question, Assist pulls the best-matching chunks into the answer. Good chunking is most of the quality:
- Aim for ~1,400 characters per chunk, split on sentence or paragraph boundaries. A chunk should read as a coherent passage on its own.
- Skip boilerplate. Navigation, footers, and cookie banners add noise, not answers. Very short fragments (under ~80 characters) aren't worth indexing.
- Keep facts with their subject. If a price or limit lives in a table, render it into the chunk as a sentence ("The free tier includes 250 questions per month") so retrieval can find it.
- One page, many chunks. Repeat the same
url/titleacross a long page's chunks — citations group by page automatically.
The Refresh Cycle
- Assist fetches your index URL hourly.
- If the
hashmatches the last ingest, nothing happens. - If it changed, the site's search index is rebuilt from the new chunks and the answer cache is cleared — stale answers can't outlive the content they came from.
Deployed something and don't want to wait? Assist → Ingestion → Re-ingest now forces a fetch immediately. The same page shows the index status: page and chunk counts, last refresh time, and the last ingest error if a fetch or parse failed.
Ingestion failures never take the assistant down — it keeps answering from the last good index until a fetch succeeds again.
Ingestion Limits
| Limit | Free | Paid |
|---|---|---|
| Pages (distinct URLs) per site | 500 | 10,000 |
| Chunks per site | 50,000 | |
Knowledge bases beyond 10,000 documents are handled by Expert Services.
Other Ingestion Sources
The index URL is the most direct path for website content. For document collections that don't live on a website, bucket ingestion points Assist at an Edge Storage bucket of markdown, text, HTML, or PDF files — with private access control built in. Crawl-based ingestion (give Assist your site URL and it follows your sitemap itself — no build step) is on the way. Every source feeds the same index and behaves identically downstream.