PHP client for the ScrapingIsNotACrime API: public data from Instagram, TikTok, YouTube, the App Store, GitHub, Hacker News, Bluesky, Twitch and Linktree.
Requires
- php: ^8.2
- php-http/discovery: ^1.19
- psr/http-client: ^1.0
- psr/http-client-implementation: ^1.0
- psr/http-factory: ^1.0
- psr/http-factory-implementation: ^1.0
- psr/http-message: ^1.1 || ^2.0
Requires (Dev)
- friendsofphp/php-cs-fixer: ^3.64
- guzzlehttp/guzzle: ^7.9 || ^8.0
- nyholm/psr7: ^1.8
- php-http/mock-client: ^1.6
- phpstan/phpstan: ^2.1
- phpunit/phpunit: ^11.5
- symfony/http-client: ^6.4 || ^7.0 || ^8.0
Suggests
- guzzlehttp/guzzle: Alternative PSR-18 client the SDK can configure
- symfony/http-client: A PSR-18 client the SDK configures with its timeout and without redirects
Provides
None
Conflicts
None
Replaces
None
This package is not auto-updated.
Last update: 2026-09-27 23:08:26 UTC
README
Official PHP SDK for the ScrapingIsNotACrime public data API — typed access to Instagram, TikTok, YouTube, the App Store, GitHub, Hacker News, Bluesky, Twitch and Linktree.
Install
composer require scrapingisnotacrime/sdk
The SDK talks over any PSR-18 HTTP client and any PSR-17 request factory. It declares those as virtual *-implementation requirements, so Composer's php-http/discovery plugin installs a PSR-18 client and PSR-17 factories for you if your project has none yet — Composer asks once whether to allow the plugin to run; answer yes. With either Symfony HttpClient or Guzzle (guzzlehttp/guzzle) present, the SDK builds its own client from whichever one it finds and configures its timeout and turns off redirects on it — but only for that client. With some other PSR-18 client already in your project, the SDK uses it as found via discovery, and that client's own timeout and redirect behavior apply instead (see Bring your own client to configure one yourself).
The explicit alternative, if you'd rather pick the implementation yourself:
composer require scrapingisnotacrime/sdk symfony/http-client nyholm/psr7
If no PSR-18 client or no PSR-17 factory can be found at all, the constructor throws \LogicException with a composer require hint.
Requires PHP 8.2+.
Quick start
use ScrapingIsNotACrime\Client; use ScrapingIsNotACrime\Exception\NotFoundException; $client = new Client(apiKey: 'sinac_...'); try { $profile = $client->instagram->profile('nasa'); echo $profile->username, ' ', $profile->followers, PHP_EOL; } catch (NotFoundException) { echo 'no such profile', PHP_EOL; }
new Client() with no arguments also reads the API key from the SCRAPINGISNOTACRIME_API_KEY environment variable. Get a key at scrapingisnotacrime.com/dashboard/api-keys. Keys start with sinac_.
Configuration
$client = new Client( apiKey: 'sinac_...', baseUrl: 'https://api.scrapingisnotacrime.com/v1', timeout: 30.0, maxRetries: 2, httpClient: null, requestFactory: null, );
| Argument | Default | Description |
|---|---|---|
apiKey |
SCRAPINGISNOTACRIME_API_KEY env var |
Your API key (sinac_…). The constructor throws \InvalidArgumentException if none is set. |
baseUrl |
https://api.scrapingisnotacrime.com/v1 |
Should be the final HTTPS URL; http is accepted for local testing; redirects are not followed, so a base URL that itself redirects fails (see Errors) — following one would forward the X-Api-Key header to whatever host it points to. |
timeout |
30.0 seconds |
Per attempt, covering connect, headers and the whole body read. Applied when the SDK builds the client itself (Symfony HttpClient or Guzzle); ignored when you pass your own httpClient. |
maxRetries |
2 |
Extra attempts for 429, 502 and network errors. 0 disables retries. |
httpClient |
a client the SDK builds itself | Bring your own PSR-18 Psr\Http\Client\ClientInterface. It is used as is: set its timeout and disable redirects yourself. |
requestFactory |
discovered via php-http/discovery |
Bring your own PSR-17 Psr\Http\Message\RequestFactoryInterface. |
Bring your own client
Pass httpClient to use a client you configure yourself — for tests, proxies or connection pooling. The SDK never modifies it, so set its own timeout and turn off redirects (the SDK does this automatically only for the client it builds itself).
Guzzle:
use GuzzleHttp\Client; use ScrapingIsNotACrime\Client as ScrapingIsNotACrimeClient; $client = new ScrapingIsNotACrimeClient( apiKey: 'sinac_...', httpClient: new Client([ 'timeout' => 30, 'allow_redirects' => false, // 'proxy' => 'http://localhost:8080', ]), );
Symfony HttpClient:
use ScrapingIsNotACrime\Client as ScrapingIsNotACrimeClient; use Symfony\Component\HttpClient\HttpClient; use Symfony\Component\HttpClient\Psr18Client; $client = new ScrapingIsNotACrimeClient( apiKey: 'sinac_...', httpClient: new Psr18Client(HttpClient::create([ 'timeout' => 30, 'max_duration' => 30, 'max_redirects' => 0, // 'proxy' => 'http://localhost:8080', ])), );
Methods
Client groups its methods under a property per platform: $client->instagram, $client->tiktok, $client->youtube, $client->appstore, $client->github, $client->hackernews, $client->bluesky, $client->twitch, $client->linktree. Every method returns the response envelope's data, decoded into a typed Types\* object; a method marked Page returns a Page instead (see Pagination). Response objects are built by the SDK; construct them with X::fromArray() if you need one in tests — constructor parameters may be reordered in minor releases.
| Namespace | Method | Arguments | Route | Page |
|---|---|---|---|---|
profile |
username |
/instagram/profile/{username} |
||
contact |
username |
/instagram/profile/{username}/contact |
||
latestPosts |
username |
/instagram/profile/{username}/timeline/latest |
||
posts |
username, count 1-50 (default 12), cursor |
/instagram/profile/{username}/timeline |
Page | |
highlights |
username |
/instagram/profile/{username}/highlights |
||
highlight |
highlightId |
/instagram/highlights/{highlightId} |
||
mediaById |
username, mediaId |
/instagram/profile/{username}/media/{mediaId} |
||
media |
shortcode |
/instagram/media/{shortcode} |
||
download |
shortcode |
/instagram/media/{shortcode}/download |
||
shortcodeToId |
shortcode |
/instagram/media/{shortcode}/id |
||
idToShortcode |
mediaId |
/instagram/media/id/{mediaId} |
||
reel |
shortcode |
/instagram/reels/{shortcode} |
||
| TikTok | profile |
username |
/tiktok/profile/{username} |
|
| TikTok | video |
videoId |
/tiktok/video/{videoId} |
|
| YouTube | videos |
handle |
/youtube/channel/{handle}/videos |
|
| App Store | search |
term, country (default "us"), limit 1-200 (default 10) |
/appstore/search |
|
| App Store | reviews |
appId, country (default "us"), page 1-10 (default 1) |
/appstore/reviews |
Page |
| GitHub | profile |
handle |
/github/profiles/{handle} |
|
| GitHub | followers |
handle, limit 1-100 (default 30), page 1-based (default 1) |
/github/profiles/{handle}/followers |
Page |
| GitHub | following |
handle, same params as followers |
/github/profiles/{handle}/following |
Page |
| GitHub | repositories |
handle, same params as followers |
/github/profiles/{handle}/repositories |
Page |
| GitHub | searchRepositories |
q (GitHub search syntax), same paging params as followers |
/github/repositories |
Page |
| GitHub | trending |
since (default daily), language, limit 1-100 (default 30) |
/github/trending/repositories |
|
| Hacker News | feed |
feed, limit 1-50 (default 20), page 0-based (default 0) |
/hackernews/feeds/{feed} |
Page |
| Hacker News | item |
id |
/hackernews/items/{id} |
|
| Hacker News | search |
q, same params as feed |
/hackernews/search |
Page |
| Hacker News | user |
username |
/hackernews/users/{username} |
|
| Hacker News | submissions |
username, same params as feed |
/hackernews/users/{username}/submissions |
Page |
| Hacker News | comments |
username, same params as feed |
/hackernews/users/{username}/comments |
Page |
| Bluesky | profile |
handle (full handle, including the domain) |
/bluesky/profiles/{handle} |
|
| Bluesky | posts |
handle, limit 1-100 (default 25), cursor |
/bluesky/profiles/{handle}/posts |
Page |
| Twitch | profile |
handle |
/twitch/profiles/{handle} |
|
| Twitch | videos |
handle, limit 1-100 (default 20) |
/twitch/profiles/{handle}/videos |
|
| Linktree | profile |
handle |
/linktree/profiles/{handle} |
since (trending's argument) accepts a GithubTrendingSince case (GithubTrendingSince::Daily/Weekly/Monthly) or its string value; feed (feed's argument) accepts a HackernewsFeed case (HackernewsFeed::Top/New/Best/Ask/Show/Job) or its string value — an unrecognized string throws \InvalidArgumentException. Every optional argument defaults to null, meaning "use the API's default". Path arguments (username, handle, shortcode, …) are validated before any request: an empty string, "." or ".." throws \InvalidArgumentException, and the value is percent-encoded the way JavaScript's encodeURIComponent would encode it.
Pagination
Methods marked Page return a ScrapingIsNotACrime\Page:
final class Page implements \IteratorAggregate { public readonly array $items; // this page's items, already decoded public readonly bool $hasMore; public readonly ?string $nextCursor; // cursor endpoints; null when there is none public readonly ?int $nextPage; // page-number endpoints; null when there is none public readonly object $data; // this page's typed page object (e.g. GithubUserPage), with fields like ->total public function next(): ?self; // fetches the following page; null when there is none public function getIterator(): \Generator; // yields items across pages, lazily }
hasMore is true exactly when next() will fetch another page (an empty final page, or an announced next page/cursor with no items, both leave it false).
Walk pages one at a time with next(). Each page fetched is one billed request, so bound the loop:
$page = $client->bluesky->posts('bsky.app', limit: 25); $fetched = 0; while ($page !== null && $fetched < 3) { foreach ($page->items as $post) { echo $post->text, PHP_EOL; } $page = $page->next(); $fetched++; }
Or iterate every item with foreach, which fetches later pages lazily. Iterating fetches every remaining page; each page is one billed request, so bound the loop:
$page = $client->github->followers('torvalds', limit: 100); $count = 0; foreach ($page as $user) { echo $user->username, PHP_EOL; if (++$count >= 250) { break; // stop early; no further pages are fetched } }
$page->data is the current page's typed page object (e.g. GithubUserPage), so fields such as ->total stay reachable.
Errors
Every failure the SDK raises is an exception under ScrapingIsNotACrime\Exception, all extending ScrapingIsNotACrimeException (itself a \RuntimeException) with ->status (the HTTP status, or null for network errors) and ->requestId (the X-Request-Id header, when the API sends one). An invalid argument — a bad path segment or an unrecognized enum string — throws \InvalidArgumentException instead, before any request is made. Numeric options such as limit, count or page are not range-checked client-side: an out-of-range value is sent to the API, which answers with a 400 and the SDK raises BadRequestException.
| Class | Status | Retried |
|---|---|---|
BadRequestException |
400 | no |
AuthenticationException |
401 | no |
QuotaExceededException |
402 | no |
NotFoundException |
404 | no |
RateLimitException |
429 | yes |
UpstreamException |
502 | yes |
ConnectionException |
network failure or per-attempt timeout; ->getPrevious() is the underlying client's exception |
yes |
ApiException |
any other status, a 2xx without the JSON envelope, a redirect, or data of an unexpected shape | no |
use ScrapingIsNotACrime\Exception\NotFoundException; use ScrapingIsNotACrime\Exception\QuotaExceededException; use ScrapingIsNotACrime\Exception\RateLimitException; use ScrapingIsNotACrime\Exception\ScrapingIsNotACrimeException; try { $profile = $client->tiktok->profile('this-user-does-not-exist-123'); echo $profile->username, PHP_EOL; } catch (NotFoundException) { echo 'no such profile', PHP_EOL; } catch (QuotaExceededException $e) { throw $e; // includes the pricing link in the message } catch (RateLimitException) { echo 'TikTok is rate limiting; already retried, try again later', PHP_EOL; } catch (ScrapingIsNotACrimeException $e) { throw $e; }
Retries
RateLimitException (429), UpstreamException (502) and ConnectionException (network failure or per-attempt timeout) are retried automatically, up to maxRetries additional attempts (default 2). These failures don't consume credits, so retrying them costs you nothing.
The delay before each retry uses the Retry-After header when the API sends one (seconds or an HTTP date); otherwise it's exponential backoff with full jitter, starting around 500 ms and doubling per attempt. Every wait is capped at 10 seconds. Set maxRetries: 0 to disable retries entirely.
Releases and changelog
Every merge to main is released automatically: the version comes from the commit messages since the last release, following Conventional Commits. The pipeline tags vX.Y.Z and publishes the GitHub Release with the notes — the tag is the release, and Packagist's GitHub hook picks up the new version, no separate publish step. The changelog is the Releases page.
Merge PRs with a merge commit or rebase so each Conventional Commit is analysed; if you squash, the PR title must be a Conventional Commit (e.g. feat: ...).
Links
- Docs: https://scrapingisnotacrime.com/docs
- Pricing: https://scrapingisnotacrime.com/#pricing
- Releases: https://github.com/ScrapingIsNotACrime/sdk-php/releases
- Go SDK: https://github.com/ScrapingIsNotACrime/sdk-go
- Python SDK: https://github.com/ScrapingIsNotACrime/sdk-python
- Node.js SDK: https://github.com/ScrapingIsNotACrime/sdk-nodejs
- MCP server: https://github.com/ScrapingIsNotACrime/mcp
- License: MIT