FAQ
Is scraping legal / ethical?
Scraping lives in a gray zone that depends on where you operate and what you scrape. Pry provides compliance tooling to help you stay on the right side:
POST /v1/compliance/check— full GDPR/CCPA compliance check on a URLPOST /v1/gdpr/*— consent recording, data deletion requests, retention policies, data portability, audit logPOST /v1/training/clean— PII/copyright stripping for training datasets
Good practice: respect robots.txt, rate-limit politely (Pry defaults to 120
RPM per IP), avoid login-walled content, and don't scrape personal data
without a lawful basis. The compliance endpoints exist precisely because
Pry's own build is GDPR-conscious.
Does Pry bypass Cloudflare?
Yes — automatically. The 15-tier fallback chain walks
direct → TLS fingerprint → cloudscraper → FlareSolverr → undetected-chromedriver →
Camoufox → Playwright → Googlebot → cookie pre-warming → behavioral biometrics →
Tor → premium proxies → archive/cache fallbacks and returns the first successful
result. Cloudflare/WAF challenges are handled by the FlareSolverr sidecar
(Docker image, health-checked by compose). You can disable it per-request with
"bypassCloudflare": false.
There is no guaranteed bypass — aggressive WAFs (e.g. DataDome on some
sites) may still win. Use POST /v1/detect-block to see which vendor is
blocking you and tune from there.
Self-hosted vs hosted — which should I use?
| Self-hosted | Hosted | |
|---|---|---|
| Cost | Free (MIT core); infrastructure is yours | Pro $49/mo, Team $199/mo |
| Privacy | All data stays on your machines | Data flows through our service |
| Ops | You run Docker/nginx/updates | Zero ops |
| Stealth module (BSL) | Free non-prod; commercial license for prod | Included |
| Best for | Privacy-sensitive teams, devs, tinkerers | Teams that don't want ops |
You can also run self-hosted and enable x402 pay-per-call to charge your own users per call — the best of both.
How does Pry scale?
Pry is a single FastAPI app + FlareSolverr + optional Redis/Postgres. To scale:
- Redis (
PRY_REDIS_URL) — shared cache backend with TTL invalidation. - Postgres (
PRY_DATABASE_URL) — async SQLAlchemy for stateful features. - Async job queue (
jobqueue.py) + WebSocket streaming for long crawls. - Horizontal: run multiple API replicas behind the load balancer; sessions and job state are the stateful parts to share (Redis-backed MCP SSE sessions are recommended for multi-worker deployments).
Resource limits in the default compose: 2 GB / 2 CPUs for Pry, 512 MB / 1 CPU for FlareSolverr. Playwright scraping is the memory-hungry part — budget accordingly.
How does Pry compare to Crawl4AI / Firecrawl / Scrapy?
| Pry | Crawl4AI | Firecrawl | Scrapy | |
|---|---|---|---|---|
| License | MIT + BSL stealth | MIT | Commercial SaaS (open-core) | BSD |
| Self-hostable | ✅ | ✅ | Limited | Framework only |
| Cloudflare bypass | ✅ 15-tier automatic | Partial | ✅ (hosted) | ❌ (DIY) |
| Browser automation | ✅ sessions, stealth, Camoufox | ⚠️ via Playwright | ⚠️ limited | ❌ (separate tools) |
| Document parsing | ✅ PDF/DOCX/OCR/CSV/JSON | ⚠️ | ✅ | ❌ |
| LLM extraction | ✅ chunked | ✅ | ✅ | ❌ |
| x402 crypto pay-per-call | ✅ native | ❌ | ❌ | ❌ |
| Compliance/GDPR tooling | ✅ built-in | ❌ | ⚠️ | ❌ |
| Commerce/CRM sync | ✅ WooCommerce, Shopify, Salesforce, HubSpot | ❌ | ❌ | ❌ |
Pry positions itself as a Firecrawl + Crawl4AI + Browserless replacement in one self-hosted app — with x402 as the differentiator: AI agents pay per call without an account, at ~$0.001–$0.10 per operation.
What auth does Pry use?
API-key and x402 gating — both fail-closed. Self-host it and set
PRY_API_KEY; every request must send Authorization: Bearer <key> (a 401
otherwise). There is no hosted account to create to use the open-source build.
Gate any endpoint with x402 so clients pay per call in crypto — the payment is
the credential. Keyless access is loopback-only by default, so an unkeyed
instance is never exposed to the internet.
Does Pry store my data?
Self-hosted: all data lives under PRY_DATA_DIR (default ~/.pry/) on your
machine — quality history, monitors, sessions, encrypted credential vault,
GDPR records. Nothing leaves your host unless you configure destinations
(webhooks, S3, GCS, SFTP) or the hosted lane.
Can I charge users with Pry?
Yes — that's the x402 pay-per-call lane. Enable PRY_X402_ENABLED=true,
point PRY_X402_PAY_TO at your wallet, and every scrape/crawl/extract earns
micropayments (USDC/USDT). See x402 Pay-per-call.
Is the stealth module free?
Free for personal, non-production, academic, and non-commercial use (BSL 1.1 Additional Use Grant) plus a 90-day evaluation. Production use requires a commercial license from [email protected]. Converts to MIT on 2029-01-01.
Where is the source?
codeberg.org/RugMunchMedia/pryscraper — mirrored to Codeberg and GitLab.
How do I report a security issue?
Email [email protected] (PGP key available on request). Response within 24 hours; critical patches within 7 days.
Next steps
- Troubleshooting — common errors and fixes
- Pricing & Monetization — the three lanes
Do I need an API key for x402 calls?
No — a verified on-chain payment is the credential. Call a paid
endpoint without paying and you get a 402 carrying a signed challenge
(amount, recipient, asset, network). Pay it with any x402 client and retry
with the payment header; the server binds the settlement to your request.
API keys still exist for teams that prefer quotas and invoicing — both
lanes work side by side.
Where do I see prices and the full catalog?
Two keyless endpoints:
GET /v1/x402/pricing— the operation price tableGET /v1/x402/catalog— every priced operation, every paid endpoint, and all 109 Apify actors bound to the MCP tool that exposes the same capability. Machine-readable, so agents can discover actor-equivalent tools and price any call without reading docs.
Can I underpay or replay a payment?
No. Server-side pricing is authoritative (client-claimed amounts are
rejected before facilitator verification), the recipient is pinned to the
configured wallet, nonces are single-use, and validAfter/validBefore
windows are enforced to the second. Tampered payloads fail signature
recovery.
How many MCP tools are there?
76 typed tools across 33 categories covering scrape, crawl, extract, monitor,
screenshots, templates, plus every data domain — generated from the same
contracts the REST API enforces, so tool docs can't drift from behavior.
Streamable HTTP + SSE at /mcp, or stdio via uvx pryscraper-mcp.