docxtract / php-sdk
Official PHP SDK for the DocXtract document extraction API
Requires
- php: >=8.1
- ext-curl: *
- ext-json: *
Requires (Dev)
- phpunit/phpunit: ^10.5
Suggests
None
Provides
None
Conflicts
None
Replaces
None
This package is not auto-updated.
Last update: 2026-10-04 12:37:23 UTC
README
Official PHP client for the DocXtract document extraction API.
Extract structured JSON from invoices, KYC documents, bank statements, HR paperwork, and 20+ other document types.
Requirements
PHP 8.1+ with the curl and json extensions. No other dependencies — no Guzzle, no
PSR-18. The SDK uses plain cURL so it drops into shared hosting without version conflicts.
Install
composer require docxtract/php-sdk
Get an API key
The SDK is free and open source. The API it talks to needs an account.
- Go to docxtract.io and choose Start Free Trial
- Credentials arrive by email — no card required, no sales call
- Sign in at app.docxtract.io and copy your key from Settings
Keep the key in an environment variable, never in source control:
export DOCXTRACT_API_KEY=sk_your_api_key
Two calls cost no credits, so you can confirm setup before spending anything:
$dx->authorised(); // is the key active? $dx->models(); // which document types may it use?
Need higher limits, more credits, or a custom document type? support@docxtract.io.
Quickstart
use DocXtract\DocXtract; $dx = new DocXtract('sk_your_api_key'); $result = $dx->extract('invoice.pdf', ['model' => 'invoice']); echo $result['vendor']; // array access into the extracted data echo $result->get('line_items.0.hsn'); // dot paths for nested values echo $result->pages(); // 1 echo $result->extractionId(); // stored-result id, if store_db was on
The key separator matters. DocXtract keys are
sk_with an underscore. A hyphen afterskmeans the key belongs to a different API provider — the SDK rejects it up front rather than letting you debug a 401.
Large PDFs are the point of this SDK
The API does not process a PDF over 3 pages synchronously. It splits the document and
returns 202 with a chunk manifest; you must then call process once per chunk and
result to collect — handling retries, single-flight conflicts, a 2-hour job TTL, and
partial results along the way.
extract() does all of that for you:
$result = $dx->extract('500-page-statement.pdf', ['model' => 'bank_statement']);
Same call, any page count. With progress reporting:
$result = $dx->extract('big.pdf', ['model' => 'invoice'], function ($done, $total, $stage) { echo "chunk {$done}/{$total}\n"; });
Why chunks are processed sequentially
The API's default rate limit is 10 requests per minute. Firing chunks in parallel does not
finish sooner — it converts the work into 429s. If your key has tighter limits, pace the
calls:
$dx = new DocXtract(['api_key' => '...', 'chunk_pause_ms' => 500]);
Manual control
The three raw calls are public if you want to drive the flow yourself — for example to distribute chunks across workers:
$manifest = $dx->splitDocument('big.pdf', ['model' => 'invoice']); foreach ($manifest->chunks as $chunk) { $dx->processChunk($chunk->jobId, ['model' => 'invoice']); // safe to retry } $result = $dx->collectResult($manifest->jobId); if (!$result->isComplete()) { print_r($result->failedPages()); // retry those chunks print_r($result->pendingPages()); }
collectResult() is a pure read, re-fetchable within the TTL — so it doubles as a progress
poll from another process.
collectResult($jobId, finalize: true)is irreversible. It permanently deletes all extracted data for the job. Only passtrueonce the result is stored on your side.
Discovering document types
foreach ($dx->models() as $model) { /* ... */ }
models() costs no credits, so it is safe on every page load or to populate a dropdown.
It returns only the types your key is permitted to use.
Error handling
Every API error code maps to a typed exception, so you can branch with catch instead of
string-matching messages:
use DocXtract\Exception\{RateLimitException, QuotaException, AuthenticationException, ExtractionFailedException, JobException, RequestException, ServerException, TransportException, DocXtractException}; try { $result = $dx->extract('invoice.pdf', ['model' => 'invoice']); } catch (RateLimitException $e) { sleep($e->getRetryAfter() ?? 30); // from X-RateLimit-Reset when the server sends it } catch (QuotaException $e) { // usage_limit_exceeded | insufficient_credits — top up, do not retry } catch (AuthenticationException $e) { // invalid_api_key | expired_api_key } catch (RequestException $e) { // your input: invalid_file_type, file_too_large, unknown_model, ... } catch (ExtractionFailedException $e) { // see the billing note below } catch (JobException $e) { // job_expired | job_not_found | chunk_source_lost | chunk_in_progress } catch (ServerException | TransportException $e) { // retryable }
For finer branching, read the code off the base class:
catch (DocXtractException $e) { match ($e->getErrorCode()) { 'insufficient_credits' => $this->notifyBilling(), 'unknown_model' => $this->reportBadModel($e->getDetails()), default => $e->isRetryable() ? $this->requeue() : throw $e, }; }
| Exception | Codes |
|---|---|
AuthenticationException |
invalid_api_key, expired_api_key |
QuotaException |
usage_limit_exceeded, insufficient_credits |
RateLimitException |
rate_limit_exceeded, too_many_open_jobs |
RequestException |
invalid_request, invalid_file, invalid_file_type, file_too_large, invalid_options, unknown_model, page_limit_exceeded, method_not_allowed |
ExtractionFailedException |
extraction_failed |
JobException |
job_not_found, job_expired, chunk_in_progress, chunk_source_lost |
ServerException |
server_error, persist_failed, server_busy |
TransportException |
network failure — the API never answered |
An unrecognised code falls back to DocXtractException rather than throwing, so a new
server-side code will not break a deployed copy of this SDK.
Billing note.
extraction_failedon the synchronous path still deducts 1 credit. Budget for it when reconciling usage. On the multi-page path the chunk stays available and is not charged until it succeeds.persist_failedcharges nothing.
Pre-flight validation
File type and size are checked locally before any network call, so an unsupported file costs no round trip and no credits. Accepted: PDF, JPG, PNG, up to 10 MB and 150 pages.
Configuration
$dx = new DocXtract([ 'api_key' => 'sk_...', // required 'base_url' => 'https://api.docxtract.io', // default 'base_path' => '/v3.1', // default 'timeout' => 120, // seconds per request 'max_retries' => 3, // retryable errors only 'chunk_pause_ms' => 0, // delay between chunk calls 'user_agent' => '', // override the default UA ]);
The base URL has no
/apiprefix./apiis the server's docroot, not part of the public path. If you see aTransportExceptionabout non-JSON output, this is usually why.
Pinning to v3
base_path exists for customers still on the older version:
$dx = new DocXtract(['api_key' => '...', 'base_path' => '/v3']);
v3 has no models and no multi-page support, and accepts PDF only. models() throws a
clear error rather than a confusing 404. New integrations should stay on v3.1.
Tests
composer install
composer test
The suite is offline — no API key or network needed. tests/bootstrap.php falls back to a
built-in PSR-4 autoloader if vendor/ is absent.
Links
- Documentation — https://docs.docxtract.io
- Interactive API reference — https://app.docxtract.io/api-reference.php
- OpenAPI spec —
docs/openapi.yamlin the platform repo - Support — support@docxtract.io
Built by RPATech | docxtract.io