Getting Started
Create a job, wait for it to be ready, run it, and retrieve structured data.
The Gluecrawl API turns a listing or directory page into structured data. This quickstart creates a job, waits for Gluecrawl to map the page, starts a run, and retrieves the resulting items.
Base URL: https://api.gluecrawl.ai/v1/
Authentication
All /v1/ endpoints require an API key passed as a Bearer token.
API access is included on every plan, free accounts included: each gets one API key. What a paid plan adds is scheduled, recurring runs — on a free account every run is one you start yourself, whether from the dashboard or through this API.
- Log in to your Gluecrawl dashboard
- Click your avatar → API Key → Generate API Key
- Copy your key — it is shown only once
Include the key in every request:
Authorization: Bearer glue_yourApiKeyHereSet it as an environment variable before using the examples below:
export GLUECRAWL_API_KEY="glue_yourApiKeyHere"The workflow
Every scrape follows the same sequence:
- Create a job with a URL and extraction input.
- Poll the job until its
statusisreadyorfailed. - Create a run for the ready job.
- Poll the run until its
statusiscompletedorfailed. - Retrieve JSON items or download CSV.
The examples use a product listing URL, but the same flow works for directories, job boards, marketplaces, and other repeating-page collections. The url doesn't need to be the exact listing page — Gluecrawl's discovery agent can navigate a site from its homepage to find the listing that matches your extraction input.
The request and response bodies below are illustrative. Replace JOB_ID and RUN_ID with the IDs returned by your own API calls, and replace the example URL and extraction goal with your target.
cURL quickstart
1. Create a job
Use a goal when you want Gluecrawl to infer the output fields from a plain-language instruction.
curl -X POST https://api.gluecrawl.ai/v1/jobs \
-H "Authorization: Bearer $GLUECRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/products",
"input": {
"type": "goal",
"value": "Extract all product names, prices, and product URLs"
},
"max_pages": 10
}'The response includes the job id and a mapping status. Keep the id; you will use it in later requests. A job progresses through in_progress, ready, or failed.
{
"id": "JOB_ID",
"status": "in_progress"
}2. Wait for the job to be ready
Request the job again until its status is ready. A ready job includes the inferred columns and protection_level.
curl "https://api.gluecrawl.ai/v1/jobs/JOB_ID" \
-H "Authorization: Bearer $GLUECRAWL_API_KEY"{
"id": "JOB_ID",
"status": "ready",
"columns": {
"listing": [
{ "name": "product_name", "type": "text" },
{ "name": "price", "type": "number" },
{ "name": "product_url", "type": "url" }
],
"detail": []
},
"protection_level": "light"
}3. Create and wait for a run
Only create a run after the job is ready.
curl -X POST "https://api.gluecrawl.ai/v1/jobs/JOB_ID/runs" \
-H "Authorization: Bearer $GLUECRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "max_pages": 10 }'The new run starts as queued. Poll GET /v1/runs/RUN_ID until its status is completed or failed.
{
"id": "RUN_ID",
"job_id": "JOB_ID",
"status": "queued",
"credits_used": null
}curl "https://api.gluecrawl.ai/v1/runs/RUN_ID" \
-H "Authorization: Bearer $GLUECRAWL_API_KEY"When the run completes, its response includes item_count, page_count, credits_used, and settled billing details.
{
"id": "RUN_ID",
"status": "completed",
"item_count": 120,
"page_count": 5,
"credits_used": 5
}4. Retrieve the data
curl "https://api.gluecrawl.ai/v1/runs/RUN_ID/items?limit=50&offset=0" \
-H "Authorization: Bearer $GLUECRAWL_API_KEY"{
"items": [
{
"data": {
"product_name": "Wireless Keyboard",
"price": 49.99,
"product_url": "https://example.com/products/wireless-keyboard"
},
"page_number": 1,
"item_index": 0
}
],
"total": 120,
"limit": 50,
"offset": 0
}For a file download instead, call the CSV export endpoint.
Node.js quickstart
This example uses Node.js with its built-in fetch. Set GLUECRAWL_API_KEY in your environment before running it.
const baseUrl = 'https://api.gluecrawl.ai/v1'
const headers = {
Authorization: `Bearer ${process.env.GLUECRAWL_API_KEY}`,
'Content-Type': 'application/json',
}
const sleep = (milliseconds) =>
new Promise((resolve) => setTimeout(resolve, milliseconds))
async function request(path, init = {}) {
const response = await fetch(`${baseUrl}${path}`, {
...init,
headers: { ...headers, ...init.headers },
})
const body = await response.json()
if (!response.ok) {
throw new Error(`${response.status}: ${body.error?.message ?? 'Request failed'}`)
}
return body
}
async function waitForStatus(path, completedStatus) {
for (;;) {
const resource = await request(path)
if (resource.status === 'failed') {
throw new Error(`Gluecrawl processing failed: ${resource.error ?? 'Unknown error'}`)
}
if (resource.status === completedStatus) return resource
await sleep(2000)
}
}
const job = await request('/jobs', {
method: 'POST',
body: JSON.stringify({
url: 'https://example.com/products',
input: {
type: 'goal',
value: 'Extract all product names, prices, and product URLs',
},
max_pages: 10,
}),
})
const readyJob = await waitForStatus(`/jobs/${job.id}`, 'ready')
const run = await request(`/jobs/${readyJob.id}/runs`, {
method: 'POST',
body: JSON.stringify({ max_pages: 10 }),
})
const completedRun = await waitForStatus(`/runs/${run.id}`, 'completed')
const items = await request(`/runs/${completedRun.id}/items?limit=50&offset=0`)
console.log(items.items)
console.log(`Used ${completedRun.credits_used} credits`)Python quickstart
Install requests, set GLUECRAWL_API_KEY, then run this script.
import os
import time
import requests
BASE_URL = "https://api.gluecrawl.ai/v1"
HEADERS = {
"Authorization": f"Bearer {os.environ['GLUECRAWL_API_KEY']}",
"Content-Type": "application/json",
}
def request(method, path, **kwargs):
response = requests.request(method, f"{BASE_URL}{path}", headers=HEADERS, **kwargs)
response.raise_for_status()
return response.json()
def wait_for_status(path, completed_status):
while True:
resource = request("GET", path)
if resource["status"] == "failed":
raise RuntimeError(f"Gluecrawl processing failed: {resource.get('error', 'Unknown error')}")
if resource["status"] == completed_status:
return resource
time.sleep(2)
job = request(
"POST",
"/jobs",
json={
"url": "https://example.com/products",
"input": {
"type": "goal",
"value": "Extract all product names, prices, and product URLs",
},
"max_pages": 10,
},
)
ready_job = wait_for_status(f"/jobs/{job['id']}", "ready")
run = request("POST", f"/jobs/{ready_job['id']}/runs", json={"max_pages": 10})
completed_run = wait_for_status(f"/runs/{run['id']}", "completed")
items = request("GET", f"/runs/{completed_run['id']}/items?limit=50&offset=0")
print(items["items"])
print(f"Used {completed_run['credits_used']} credits")Polling, retries, and credits
Polling
- Poll a job until it is
readyorfailed; only a ready job can create a run. - Poll a run until it is
completedorfailed. Items can appear progressively while a run is in progress, but a completed run contains final counts and settled billing. - The two-second interval in the examples is illustrative. Keep the total requests for an API key within the 60 requests-per-minute limit, especially when you run several workflows at once.
Retries
- For
429 rate_limited, wait for the response'sRetry-Aftervalue before making another request. - Do not retry
401,402,403,404,409, or422blindly. Fix the API key, credits, plan, resource/state, or request body first. See Error Codes. - If your client loses a response to a write request, first check the relevant job or run list before creating another resource. This avoids treating an unknown outcome as a safe duplicate.
Credits
- Creating a job costs 10 credits. The charge is refunded automatically if job processing fails.
- Runs are charged after completion for the work actually performed. Cost depends on the selected protection level, listing pages, and any detail-page visits.
- Set
max_pagesdeliberately, then use a completed run'scredits_usedandbillingfields to understand its final cost. See Credits for the rate table.
Skip the polling loop with n8n
If you would rather not write the loop above, Gluecrawl is available as an n8n node that is verified by n8n, so it installs on n8n Cloud and on self-hosted instances alike. Its trigger node wakes a workflow when a job is ready or a run completes, which removes the polling entirely. See Use Gluecrawl in n8n.