AI Crawler Tracking
Your existing Metrivo tracker automatically records supported crawlers that execute it; no second browser script is needed. Most AI and search crawlers do not execute JavaScript. Optional server-log forwarding adds observations of requests that reach your collector in the AI Crawlers dashboard.
Tracked crawlers
Metrivo classifies these user agents. Any other User-Agent is dropped on ingest:
OAI-SearchBotChatGPT-UserGPTBotClaude-SearchBotClaude-UserClaudeBotPerplexityBotGooglebotBingbot
1. Create an API key with crawlers:write scope
In Settings → API Keys, create a new key with the crawlers:write scope. Scope the key to a single website if you can — that prevents accidental cross-website writes.
The key is workspace- or website-scoped at validation time; the endpoint rejects writes to any website outside the key's workspace.
2. POST request batches to /api/crawlers/log
Send up to 500 items per request. Each item is a single crawler hit.
curl -X POST https://www.metrivo.co/api/crawlers/log \
-H "Authorization: Bearer <API_KEY_WITH_crawlers:write_SCOPE>" \
-H "Content-Type: application/json" \
-d '{
"websiteId": "<WEBSITE_UUID>",
"items": [
{
"userAgent": "Mozilla/5.0 (compatible; GPTBot/1.0; +https://openai.com/gptbot)",
"path": "/pricing",
"statusCode": 200,
"accessState": "allowed",
"occurredAt": "2026-05-21T14:32:18.000Z"
}
]
}'Required fields per item
userAgent— raw User-Agent header (max 2048 chars). Items not matching the tracked-bot list are silently dropped.
Optional fields
path— pathname only. If you sendurlinstead, Metrivo extracts the path.statusCode— HTTP status code returned to the crawler. NULL means "Unknown" in the dashboard.accessState— "allowed" or "blocked". If your origin returned 403 or the crawler was blocked by your WAF, set "blocked". Otherwise omit it — Metrivo will show "Unknown" rather than guess.occurredAt— ISO-8601 timestamp. Defaults to ingestion time.
3. Privacy and tenant safety
- Strip IPs, cookies, query strings, and any user-identifying headers before forwarding. Metrivo stores only the User-Agent, path, status code, access state, and timestamp.
- The endpoint enforces that the API key's workspace owns the target
websiteId. Cross-workspace writes are rejected with 403. - If the API key is scoped to a single website, the
websiteIdfield is optional — the endpoint uses the key's scope.
4. Automatic collection with Cloudflare Workers
- Download the Worker and create a Cloudflare Worker using its module code. If your website already has a Worker, merge this forwarding logic into its existing handler.
- Add
METRIVO_API_KEYas an encrypted Worker secret containing yourcrawlers:writekey. AddMETRIVO_WEBSITE_IDas a variable using the website ID shown in the crawler explorer. Never put the key in browser code. - Attach a Worker route matching your own proxied website, such as
example.com/*, so requests pass through it to your existing origin. Do not attach it to metrivo.co or use a Worker custom domain as a replacement origin. - Deploy the Worker. After a real supported crawler request, check the last server-log receipt in AI crawler traffic. A saved key alone does not confirm collection.
The forwarder sends one event per observed bot request in the background using Cloudflare waitUntil. It leaves visitor responses intact if delivery fails and reports failures in Worker logs. Delivery is best-effort with no automatic retries; use a durable log pipeline for high-volume or loss-sensitive reporting. Do not forward the same request through multiple collectors: ingestion does not deduplicate repeated submissions.
Only requests reaching this Worker are observable. Blocks applied before it runs require a separate WAF log feed. HTTP status is recorded; access remains unknown unless a 401 or 403 provides evidence of denial. Historic requests are not recovered automatically.
Other server-log pipelines
Forward from an Nginx / Cloudflare / Vercel access log, or from a small middleware in your app:
# Inside your access log post-processor (or a small forwarder):
# 1. Filter for AI/search bot user agents from your access log.
# 2. POST batches of up to 500 items to /api/crawlers/log.
# 3. Drop or hash any IP / cookie data before forwarding.5. What Metrivo will not claim
- If you omit
statusCodeoraccessState, the dashboard shows "Unknown". We never fabricate a value. - We do not infer that a missing log entry means a crawler did not visit — it might mean the log was not forwarded.
- Bot names are user-agent matches, not verified identities. Googlebot requests do not identify Gemini use. See Google crawler documentation.
- The explorer aggregates the complete selected time window, paginates pages and requests, and exports only the visible table rows. Filters apply to all reported metrics; selector counts show the whole date range.
- We do not deduplicate identical hits across multiple POSTs. Send each hit exactly once.
