AgentStack
SKILL verified MIT Self-run

Scraperapi Php Sdk

skill-scraperapi-scraperapi-skills-scraperapi-php-sdk · by scraperapi

>

No reviews yet
0 installs
17 views
0.0% view→install

Install

$ agentstack add skill-scraperapi-scraperapi-skills-scraperapi-php-sdk

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Scraperapi Php Sdk? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

ScraperAPI — PHP SDK Best Practices

Requires: PHP 7.0+, Composer, composer require scraperapi/sdk, SCRAPERAPI_API_KEY environment variable.

Setup

raw_body
$html = $client->get("https://example.com/")->raw_body;
echo $html;

// With a single parameter
$html = $client->get("https://example.com/", ["render" => true])->raw_body;

// With multiple parameters
$html = $client->get(
    "https://example.com/",
    [
        "render"       => true,
        "country_code" => "us",
    ]
)->raw_body;

Parameters are passed as an associative array. ->raw_body extracts the HTML string from the response object.

Decision Guide

| Situation | Approach | |-----------|---------| | Single URL, synchronous | $client->get($url, $params)->raw_body | | Page loads content via JavaScript | Add "render" => true | | Site blocks datacenter proxies | Add "premium" => true | | Toughest anti-bot protection | Add "ultra_premium" => true | | Multi-step / paginated flow on same domain | Use "session_number" | | POST a form or JSON body to target | $client->post($url, $options)->raw_body | | 20+ URLs or batch jobs | Use async endpoint via cURL or Guzzle | | Supported platform (Amazon, Google, etc.) | Use structured data endpoint directly |

Parameter Reference

Rendering

// Render JavaScript before returning HTML
// Use when: page is a React/Vue/Angular SPA, or scrape returns empty/partial content
// Cost: +10 credits
$html = $client->get("https://spa-site.com/", ["render" => true])->raw_body;

// Wait for a specific DOM element (requires render: true)
$html = $client->get("https://spa-site.com/", [
    "render"            => true,
    "wait_for_selector" => ".product-list",
])->raw_body;

// Screenshot (auto-enables rendering)
$html = $client->get("https://example.com/", ["screenshot" => true])->raw_body;

Start without render. Add it only when the response is missing expected content — it increases cost and latency.

Proxies and Geotargeting

// Route through a country-specific proxy — no extra credit cost
$html = $client->get("https://example.com/", ["country_code" => "de"])->raw_body;

// Premium residential/mobile IPs — for sites that block datacenter proxies
// Cost: 10 credits (25 with render)
$html = $client->get("https://hard-site.com/", ["premium" => true])->raw_body;

// Ultra-premium — for the toughest anti-bot protections
// Cost: 30 credits (75 with render)
// Note: incompatible with custom headers — keep_headers is ignored
$html = $client->get("https://hardest-site.com/", ["ultra_premium" => true])->raw_body;

premium and ultra_premium are mutually exclusive — never set both. Escalation order: standard (1 cr) → render (10 cr) → premium (10 cr) → ultra_premium (30 cr).

Sessions (Sticky Proxy)

// Reuse the same proxy IP across requests — useful for pagination and multi-step flows
// Sessions expire 15 minutes after last use; any integer is a valid session ID
$html1 = $client->get("https://example.com/page1", ["session_number" => 42])->raw_body;
$html2 = $client->get("https://example.com/page2", ["session_number" => 42])->raw_body;

Headers and Device Type

// Forward custom headers to the target site
// Note: keep_headers is ignored when ultra_premium is true
$html = $client->get("https://example.com/", [
    "keep_headers" => true,
])->raw_body;
// Pass headers in the request options array alongside ScraperAPI params

// Emulate a mobile or desktop browser user-agent
$html = $client->get("https://example.com/", ["device_type" => "mobile"])->raw_body;

Autoparse and Response Format

// Return structured JSON instead of HTML for supported sites (Amazon, Google, etc.)
$json = $client->get("https://amazon.com/dp/B09V3KXJPB", ["autoparse" => true])->raw_body;
$data = json_decode($json, true);

// Markdown output — useful for text pipelines
$md = $client->get("https://docs.example.com/", ["output_format" => "markdown"])->raw_body;

POST Requests

// POST a JSON body to the target site through ScraperAPI's proxy
$options = [
    "body"    => json_encode(["key" => "value"]),
    "headers" => ["Content-Type" => "application/json"],
];
$result = $client->post("https://example.com/api", $options)->raw_body;

Escalation Ladder

Always start with the cheapest option and escalate only when blocked.

function scrapeWithEscalation(Client $client, string $url): ?string
{
    $tiers = [
        [],
        ["render"       => true],
        ["premium"      => true],
        ["premium"      => true, "render" => true],
        ["ultra_premium" => true],
    ];

    foreach ($tiers as $params) {
        $html = $client->get($url, $params)->raw_body;
        if ($html && stripos($html, 'get()` call blocks until the response (up to 70 seconds). For 20+ URLs, use the async REST endpoint.

```php
$apiKey = getenv('SCRAPERAPI_API_KEY');

function submitJob(string $url, array $apiParams = []): array
{
    global $apiKey;
    $ch = curl_init('https://async.scraperapi.com/jobs');
    curl_setopt_array($ch, [
        CURLOPT_POST           => true,
        CURLOPT_POSTFIELDS     => json_encode(['apiKey' => $apiKey, 'url' => $url, 'apiParams' => $apiParams]),
        CURLOPT_HTTPHEADER     => ['Content-Type: application/json'],
        CURLOPT_RETURNTRANSFER => true,
    ]);
    $response = curl_exec($ch);
    curl_close($ch);
    return json_decode($response, true); // ["id" => "...", "statusUrl" => "..."]
}

function pollJob(array $job, int $maxWait = 120, int $interval = 5): string
{
    $deadline = time() + $maxWait;
    while (time()  $apiKey], $params));
    $url   = "https://api.scraperapi.com/structured/{$vertical}?{$query}";
    $body  = file_get_contents($url);
    if ($body === false) throw new \RuntimeException("Request failed for {$vertical}");
    return json_decode($body, true);
}

// Google SERP
$results = structuredGet('google/search', ['query' => 'PHP web scraping']);

// Amazon product details
$product = structuredGet('amazon/product', ['asin' => 'B09V3KXJPB']);

// Walmart search
$items = structuredGet('walmart/search', ['query' => 'standing desk', 'tld' => 'com']);

See structured data docs for all verticals and required fields.

Error Handling

function safeScrape(Client $client, string $url, array $params = []): ?string
{
    try {
        return $client->get($url, $params)->raw_body;
    } catch (\Exception $e) {
        $status = method_exists($e, 'getCode') ? (int) $e->getCode() : 0;
        switch ($status) {
            case 401: throw new \RuntimeException('Invalid API key — check SCRAPERAPI_API_KEY');
            case 403: throw new \RuntimeException('Blocked or out of credits — try premium or ultra_premium');
            case 429: throw new \RuntimeException('Rate limit — reduce concurrency or switch to async');
            case 500:
            case 503: throw new \RuntimeException('Transient error — retry with exponential backoff');
            default:  throw $e;
        }
    }
}

Status code reference: 200 success, 401 bad key, 403 blocked/no credits, 404 target not found, 429 rate limit, 500/503 transient (not charged — safe to retry).

Also see retry docs.

Credit Cost Reference

| Request type | Credits | |---|---| | Standard | 1 | | "render" => true | 10 | | "premium" => true | 10 | | "premium" => true, "render" => true | 25 | | "ultra_premium" => true | 30 | | "ultra_premium" => true, "render" => true | 75 |

Add "max_cost" => N to any request to cap credit spend — returns 403 if the request would cost more than N credits.

Documentation

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.