webcrawlerapi / sdk
A PHP SDK for WebCrawler API - turn website into data
Requires
- php: >=8.1
- ext-json: *
- guzzlehttp/guzzle: ^7.0
Requires (Dev)
- guzzlehttp/psr7: ^2.0
- phpunit/phpunit: ^10.0
Suggests
None
Provides
None
Conflicts
None
Replaces
None
README
A PHP SDK for interacting with the WebCrawlerAPI - a powerful web crawling and scraping service.
In order to use the API you have to get an API key from WebCrawlerAPI
Read documentation at WebCrawlerAPI Docs for more information.
Requirements
- PHP 8.0 or higher
- Composer
ext-jsonPHP extension- Guzzle HTTP Client 7.0 or higher
Installation
You can install the package via composer:
composer require webcrawlerapi/sdk
Usage
Scrape a single page and get its content as markdown:
use WebCrawlerAPI\Models\ScrapeRequest; use WebCrawlerAPI\Models\ScrapeResponseError; use WebCrawlerAPI\WebCrawlerAPI; $client = new WebCrawlerAPI('your_api_key'); $result = $client->scrape(new ScrapeRequest( url: 'https://example.com', outputFormats: ['markdown'], )); if ($result instanceof ScrapeResponseError) { echo "Scrape failed: {$result->errorCode} - {$result->errorMessage}\n"; } else { echo $result->pageTitle . "\n"; echo $result->markdown . "\n"; }
scrape() blocks until the result is ready. For non-blocking use there are also scrapeAsync() and getScrape().
Scrape parameters
ScrapeRequest accepts:
url(required): The page to scrape.outputFormats(optional): Array ofmarkdown,cleaned,html,links.prompt(optional): AI extraction prompt. Result is returned instructuredData.responseSchema(optional): JSON schema for the structured output ofprompt.cleanSelectors(optional): CSS selectors of elements to remove.mainContentOnly(optional): Return only the main content of the page.respectRobotsTxt(optional): Respect the site's robots.txt.maxAge(optional): Maximum age of a cached result in seconds. Use0to always fetch fresh.webhookUrl(optional): URL that receives a POST request once the scrape is done.
Scrape response
ScrapeResponse has success, status, markdown, cleanedContent, rawContent, links, structuredData, pageTitle and pageStatusCode. API errors are returned as a ScrapeResponseError with errorCode and errorMessage.
Crawling
To crawl a whole site, use crawl() or crawlAsync(). See the crawling guide.
Testing
Running Tests
-
Install dependencies:
composer install
-
Run unit tests:
vendor/bin/phpunit tests/Unit --testdox
-
Run integration tests (optional, requires API key):
export WEBCRAWLER_API_KEY="your-api-key" vendor/bin/phpunit tests/Integration --testdox
Or use the test runner script: ./run-tests.sh
License
MIT License