Search by

A PHP SDK for WebCrawler API - turn website into data

1.2.0 2026-09-21 17:27 UTC

This package is auto-updated.

Last update: 2026-09-21 17:45:53 UTC


README

Latest Version on Packagist Total Downloads License

A PHP SDK for interacting with the WebCrawlerAPI - a powerful web crawling and scraping service.

In order to use the API you have to get an API key from WebCrawlerAPI

Read documentation at WebCrawlerAPI Docs for more information.

Requirements

  • PHP 8.0 or higher
  • Composer
  • ext-json PHP extension
  • Guzzle HTTP Client 7.0 or higher

Installation

You can install the package via composer:

composer require webcrawlerapi/sdk

Usage

Scrape a single page and get its content as markdown:

use WebCrawlerAPI\Models\ScrapeRequest;
use WebCrawlerAPI\Models\ScrapeResponseError;
use WebCrawlerAPI\WebCrawlerAPI;

$client = new WebCrawlerAPI('your_api_key');

$result = $client->scrape(new ScrapeRequest(
    url: 'https://example.com',
    outputFormats: ['markdown'],
));

if ($result instanceof ScrapeResponseError) {
    echo "Scrape failed: {$result->errorCode} - {$result->errorMessage}\n";
} else {
    echo $result->pageTitle . "\n";
    echo $result->markdown . "\n";
}

scrape() blocks until the result is ready. For non-blocking use there are also scrapeAsync() and getScrape().

Scrape parameters

ScrapeRequest accepts:

  • url (required): The page to scrape.
  • outputFormats (optional): Array of markdown, cleaned, html, links.
  • prompt (optional): AI extraction prompt. Result is returned in structuredData.
  • responseSchema (optional): JSON schema for the structured output of prompt.
  • cleanSelectors (optional): CSS selectors of elements to remove.
  • mainContentOnly (optional): Return only the main content of the page.
  • respectRobotsTxt (optional): Respect the site's robots.txt.
  • maxAge (optional): Maximum age of a cached result in seconds. Use 0 to always fetch fresh.
  • webhookUrl (optional): URL that receives a POST request once the scrape is done.

Scrape response

ScrapeResponse has success, status, markdown, cleanedContent, rawContent, links, structuredData, pageTitle and pageStatusCode. API errors are returned as a ScrapeResponseError with errorCode and errorMessage.

Crawling

To crawl a whole site, use crawl() or crawlAsync(). See the crawling guide.

Testing

Running Tests

  1. Install dependencies:

    composer install
  2. Run unit tests:

    vendor/bin/phpunit tests/Unit --testdox
  3. Run integration tests (optional, requires API key):

    export WEBCRAWLER_API_KEY="your-api-key"
    vendor/bin/phpunit tests/Integration --testdox

Or use the test runner script: ./run-tests.sh

License

MIT License