API Overview
Pry exposes a FastAPI JSON API. All endpoints are documented in the bundled
OpenAPI spec (openapi.json in the repo root) and served live at
/docs (Swagger UI) when the server is running.
Base URL
| Environment | Base URL |
|---|---|
| Local (bare metal) | http://localhost:8002 |
| Docker (host) | https://api.pryscraper.com |
| Configurable | PRY_URL env var |
All request/response bodies are JSON. There are 251 registered paths across 64 tag groups.
Authentication
Authentication is enforced by the PryHttpMiddleware and request_authorized.
The policy is fail-closed:
PRY_API_KEYset → EVERY request (loopback or remote) must sendAuthorization: Bearer <key>(or an rmi JWT). Requests without a valid credential get401 Unauthorized.PRY_API_KEYunset → the API is only reachable from the loopback interface (127.0.0.1/::1). Every non-loopback request is rejected with401, so a keyless instance is never exposed to the internet (e.g. when the Docker port mapping is public).
curl -X POST https://api.pryscraper.com/v1/scrape \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <your-api-key>" \
-d '{"url": "https://example.com"}'
Public paths that skip auth: /health, /live, /ready (liveness/readiness
probes).
Proxy/Tor configuration endpoints (/v1/proxy/configure, POST /v1/config,
/v1/config/profile/tor) are additionally guarded: remote clients that cannot
present the key are rejected.
Rate limits
- Token bucket per IP, default 120 requests/minute (
PRY_RATE_LIMIT_RPM). - Exceeding the limit returns
429 Too Many Requestswith:Retry-Afterresponse header (seconds)x-ratelimit-limit,x-ratelimit-remaining,x-ratelimit-resetheaders- A body containing
retry_after
{
"type": "/errors/rate_limit_exceeded",
"title": "Too Many Requests",
"status": 429,
"detail": "Rate limit exceeded",
"request_id": "a1b2c3d4e5f6",
"instance": "/v1/scrape",
"timestamp": "2026-08-16T12:00:00.000000+00:00",
"code": "rate_limit_exceeded",
"retry_after": 1
}
Error format
Errors follow RFC 7807 Problem Details, returned by the global exception handler:
{
"type": "/errors/unauthorized",
"title": "Unauthorized",
"status": 401,
"detail": "Invalid or missing API key",
"request_id": "a1b2c3d4e5f6",
"instance": "https://api.pryscraper.com/v1/scrape",
"timestamp": "2026-08-16T12:00:00.000000+00:00",
"code": "unauthorized"
}
| Field | Meaning |
|---|---|
type | Error type URI (/errors/<slug>, or about:blank for 500s) |
title | Human-readable title |
status | HTTP status code |
detail | Human-readable detail |
request_id | Correlation ID (also echoed in the x-request-id response header) |
instance | The request URL that produced the error |
timestamp | ISO 8601 UTC timestamp |
code | Machine-readable slug of the error |
Common status codes:
| Code | Meaning |
|---|---|
400 | Bad request (invalid input) |
401 | Unauthorized — missing/invalid API key, or remote client with no key set |
403 | Forbidden |
404 | Not found (unknown path) |
409 | Conflict |
422 | Validation error (Pydantic) |
429 | Rate limit exceeded |
500 | Internal server error (detail is hidden, use request_id) |
502 / 503 | Upstream / service unavailable |
402 | Payment required — only when x402 gating is enabled (see x402 Pay-per-call) |
Health & monitoring endpoints
| Endpoint | Purpose |
|---|---|
GET /health | Service health + cache stats + active sessions |
GET /live | Liveness probe |
GET /ready | Readiness probe |
GET /metrics | Prometheus metrics |
GET /v0/stats | Basic stats |
Request IDs
Every request gets a request_id. Send your own with the x-request-id
header to correlate logs end-to-end; otherwise a UUID is generated.
API groups
The API is organized into tag groups — the most relevant for day-to-day use:
| Tag | Key endpoints |
|---|---|
| Health | /health, /live, /ready |
| Scraping | /v1/scrape, /v1/crawl, /v1/map, /v1/batch, /v1/ultimate-scrape, /v1/detect-block |
| Extraction | /v1/extract, /v1/extract/css, /v1/extract/llm, /v1/parse, /v1/shadow-dom, /v1/schema |
| Automation | /v1/automate, /v1/screenshot, /v1/session/*, /v1/capture/* |
| x402 | /v1/x402/pricing, /v1/x402/pay, /v1/x402/verify, /v1/x402/payment, /v1/x402/require-payment |
| Batch | /v1/batch, /v1/batch-file |
| Monitoring | /v1/watch, /v1/monitor, /v1/freshness/*, /v1/diff |
| Analysis | /v1/vision, /v1/summarize, /v1/categorize, /v1/compare |
| Sessions | /v1/session/create, /v1/session/save, /v1/session/restore, /v1/session/destroy, /v1/sessions |
| MCP | /mcp/tools, /mcp/call (see MCP Integration for the supported transports) |
All routes by data domain
Every domain route below is a real endpoint from routers/*.py — grouped by
the Data Domains sections. Core platform routes
(scraping, extraction, monitoring, GDPR, pipelines, …) are listed in the tag
table above; the complete machine-readable surface ships as openapi.json
in the repo root.
Crypto Market — 3 routes
See Crypto Market.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/crypto/token-launches | New tokens / pairs / trending |
| GET | /v1/crypto/cex-listings | CEX listing / delisting signals |
| GET | /v1/crypto/token-security/{chain}/{address} | Rug / honeypot security scan |
Marketplace — 7 routes
See Marketplace.
| Method | Path | Summary |
|---|---|---|
| POST | /v1/actors/create | Create a marketplace actor |
| GET | /v1/actors | List actors (filter by visibility / tag) |
| POST | /v1/actors/{actor_id}/run | Run an actor |
| GET | /v1/marketplace/amazon/product/{asin} | Amazon product by ASIN |
| GET | /v1/marketplace/amazon/search | Amazon keyword search |
| GET | /v1/marketplace/walmart/product/{product_id} | Walmart product |
| GET | /v1/marketplace/tiktok/trending | TikTok Shop trending |
Legal & Public Records — 5 routes
| Method | Path | Summary |
|---|---|---|
| GET | /v1/legal/dockets/search | CourtListener docket search |
| GET | /v1/legal/dockets/{docket_id} | Single docket lookup |
| GET | /v1/legal/sec/search | SEC EDGAR filing search |
| GET | /v1/legal/sec/insider/{ticker} | Form 4 insider trades |
| GET | /v1/legal/sanctions/{name} | OFAC SDN name screening |
Real Estate — 3 routes
See Real Estate.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/realestate/listing/{zillow_id} | Zillow listing lookup |
| GET | /v1/realestate/search | Listings by location / status |
| GET | /v1/realestate/property/{county}/{parcel_id} | County property record |
Jobs — 6 routes
See Jobs.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/job/{job_id} | Async job status + result |
| GET | /v1/jobs | List jobs in a Pryfile |
| GET | /v1/jobs/search | Cross-board keyword search |
| GET | /v1/jobs/salary/{company} | Levels.fyi salary bands |
| GET | /v1/jobs/{board}/{company} | ATS job listings |
| GET | /v1/jobs/{board}/job/{job_id} | One ATS job's full detail |
Travel — 3 routes
See Travel.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/travel/flights/grid | Cheapest flight per day over a range |
| GET | /v1/travel/flights/search | One-way flight options |
| GET | /v1/travel/hotels/rates | Per-night hotel rates |
Attention & SEO — 7 routes
See Attention & SEO.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/attention/serp | Google SERP (organic + AI Overview + PAA) |
| GET | /v1/attention/trends | Google Trends interest |
| GET | /v1/attention/ads | Meta Ad Library search |
| POST | /v1/seo/analyze | SEO element analysis |
| POST | /v1/seo/track | SEO change tracking |
| POST | /v1/seo/keywords | Keyword presence / density |
| POST | /v1/seo | Legacy SEO analysis (extraction router) |
Local Business — 3 routes
See Local Business.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/local/places/search | Google Maps business search |
| GET | /v1/local/places/{place_id} | Place detail + popular times |
| GET | /v1/local/reviews/{place_id} | Google / Yelp review history |
Document Parsing — 6 routes
See Document Parsing.
| Method | Path | Summary |
|---|---|---|
| POST | /v1/parse | Parse PDF/DOCX/OCR/CSV/JSON |
| POST | /v1/markdown | Markdown with content filtering |
| POST | /v1/shadow-dom | Shadow DOM extraction |
| POST | /v1/pdf/extract | PDF table extraction |
| POST | /v1/ocr/extract | Image OCR |
| POST | /v1/extract-table | HTML table extraction |
Browser Automation — 15 routes
See Browser Automation.
| Method | Path | Summary |
|---|---|---|
| POST | /v1/automate | Step-based browser automation |
| POST | /v1/screenshot | Screenshot → base64 PNG |
| POST | /v1/session/create | Create persistent session |
| POST | /v1/session/destroy | Destroy session |
| GET | /v1/sessions | List sessions |
| POST | /v1/session/save | Save session state to disk |
| POST | /v1/session/restore | Restore saved session |
| POST | /v1/record/start | Start recording browser actions |
| POST | /v1/record/step | Record an action step |
| POST | /v1/record/export | Export recording as script |
| POST | /v1/record/clear | Clear recorded actions |
| POST | /v1/capture/lazy | Lazy-load / infinite-scroll capture |
| POST | /v1/capture/network | Network / hidden-API capture |
| POST | /v1/ws/scrape | WebSocket data capture |
| POST | /v1/sse/scrape | Server-Sent Events capture |
Sports Data — 13 routes
See Sports Betting.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/sports/odds | Live betting odds across bookmakers |
| GET | /v1/sports/odds/{event_id} | Multi-sportsbook odds comparison |
| GET | /v1/sports/stats/{player} | Player stats (ESPN) |
| GET | /v1/sports/scores | Final / live scores for a day |
| GET | /v1/sports/arbitrage | Guaranteed-profit arbitrage across bookmakers |
| GET | /v1/sports/plus-ev | Positive-EV betting vs a sharp book |
| GET | /v1/sports/props | Player prop betting lines |
| GET | /v1/sports/injuries | Injury report for a sport or team |
| GET | /v1/sports/live | Live in-play odds |
| GET | /v1/sports/historical | Historical betting odds for past events |
| GET | /v1/sports/betting-percentages | Public betting percentages (bets & money) |
| GET | /v1/sports/sharp-consensus | Sharp consensus lines (CLV-aware) |
| GET | /v1/sports/alerts | Money-printer alerts (line moves, steam, CLV) |
Healthcare Data — 4 routes
See Healthcare.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/health/trials/search | Clinical trial keyword search |
| GET | /v1/health/trials/{nct_id} | One clinical trial by NCT id |
| GET | /v1/health/drug/{drug} | US drug price (NADAC + GoodRx fallback) |
| GET | /v1/health/insurance/quotes | Insurance quote feasibility + public rate tables |
Commerce Data — 5 routes
See Commerce Long-Tail.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/commerce/search | Product search across eBay/Etsy/AliExpress/Shopify |
| GET | /v1/commerce/platforms | List supported marketplaces / platforms |
| GET | /v1/commerce/sync | Commerce sync status / targets |
| GET | /v1/commerce/{marketplace}/{product_id} | Single product by id |
| GET | /v1/social/{platform}/{handle} | Public follower metrics + fake-follower estimate |
Government Data — 3 routes
See Government Data.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/gov/procurement/search | Federal procurement opportunities (SAM.gov) |
| GET | /v1/gov/procurement/{notice_id} | Single procurement notice |
| GET | /v1/gov/recalls/search | CPSC / FDA product recalls |
IP Data — 4 routes
See IP Data.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/ip/patents/search | US patent keyword search (Google Patents) |
| GET | /v1/ip/patents/{patent_id} | Single patent detail |
| GET | /v1/ip/trademarks/search | US trademark search (USPTO) |
| GET | /v1/ip/companies/search | State company registry search |
Climate Data — 2 routes
| Method | Path | Summary |
|---|---|---|
| GET | /v1/climate/weather | NOAA NWS current weather + short-range forecast |
| GET | /v1/climate/agri | USDA NASS crop prices / agri conditions |
Research Data — 2 routes
See Research & Academic.
| Method | Path | Summary |
|---|---|---|
| GET | /v1/research/papers | Academic paper / preprint search (OpenAlex + Crossref) |
| GET | /v1/research/domain | Domain intelligence report (RDAP WHOIS + crt.sh) |
Total: 91 domain routes across the 17 data domains (the full API has 251 registered paths).
Next steps
- Scraping API — scrape, crawl, batch, map, ultimate-scrape
- Extraction API — CSS/LLM extraction, parse, shadow DOM
- Automation API — automate, sessions, capture