Search by

docxtract / php-sdk

RPATech

Official PHP SDK for the DocXtract document extraction API

v1.0.0 2026-09-05 07:58 UTC

This package is not auto-updated.

Last update: 2026-10-04 12:37:23 UTC


README

Official PHP client for the DocXtract document extraction API.

Extract structured JSON from invoices, KYC documents, bank statements, HR paperwork, and 20+ other document types.

Requirements

PHP 8.1+ with the curl and json extensions. No other dependencies — no Guzzle, no PSR-18. The SDK uses plain cURL so it drops into shared hosting without version conflicts.

Install

composer require docxtract/php-sdk

Get an API key

The SDK is free and open source. The API it talks to needs an account.

  1. Go to docxtract.io and choose Start Free Trial
  2. Credentials arrive by email — no card required, no sales call
  3. Sign in at app.docxtract.io and copy your key from Settings

Keep the key in an environment variable, never in source control:

export DOCXTRACT_API_KEY=sk_your_api_key

Two calls cost no credits, so you can confirm setup before spending anything:

$dx->authorised();   // is the key active?
$dx->models();       // which document types may it use?

Need higher limits, more credits, or a custom document type? support@docxtract.io.

Quickstart

use DocXtract\DocXtract;

$dx = new DocXtract('sk_your_api_key');

$result = $dx->extract('invoice.pdf', ['model' => 'invoice']);

echo $result['vendor'];              // array access into the extracted data
echo $result->get('line_items.0.hsn');  // dot paths for nested values
echo $result->pages();               // 1
echo $result->extractionId();        // stored-result id, if store_db was on

The key separator matters. DocXtract keys are sk_ with an underscore. A hyphen after sk means the key belongs to a different API provider — the SDK rejects it up front rather than letting you debug a 401.

Large PDFs are the point of this SDK

The API does not process a PDF over 3 pages synchronously. It splits the document and returns 202 with a chunk manifest; you must then call process once per chunk and result to collect — handling retries, single-flight conflicts, a 2-hour job TTL, and partial results along the way.

extract() does all of that for you:

$result = $dx->extract('500-page-statement.pdf', ['model' => 'bank_statement']);

Same call, any page count. With progress reporting:

$result = $dx->extract('big.pdf', ['model' => 'invoice'], function ($done, $total, $stage) {
    echo "chunk {$done}/{$total}\n";
});

Why chunks are processed sequentially

The API's default rate limit is 10 requests per minute. Firing chunks in parallel does not finish sooner — it converts the work into 429s. If your key has tighter limits, pace the calls:

$dx = new DocXtract(['api_key' => '...', 'chunk_pause_ms' => 500]);

Manual control

The three raw calls are public if you want to drive the flow yourself — for example to distribute chunks across workers:

$manifest = $dx->splitDocument('big.pdf', ['model' => 'invoice']);

foreach ($manifest->chunks as $chunk) {
    $dx->processChunk($chunk->jobId, ['model' => 'invoice']);  // safe to retry
}

$result = $dx->collectResult($manifest->jobId);

if (!$result->isComplete()) {
    print_r($result->failedPages());   // retry those chunks
    print_r($result->pendingPages());
}

collectResult() is a pure read, re-fetchable within the TTL — so it doubles as a progress poll from another process.

collectResult($jobId, finalize: true) is irreversible. It permanently deletes all extracted data for the job. Only pass true once the result is stored on your side.

Discovering document types

foreach ($dx->models() as $model) { /* ... */ }

models() costs no credits, so it is safe on every page load or to populate a dropdown. It returns only the types your key is permitted to use.

Error handling

Every API error code maps to a typed exception, so you can branch with catch instead of string-matching messages:

use DocXtract\Exception\{RateLimitException, QuotaException, AuthenticationException,
    ExtractionFailedException, JobException, RequestException, ServerException,
    TransportException, DocXtractException};

try {
    $result = $dx->extract('invoice.pdf', ['model' => 'invoice']);
} catch (RateLimitException $e) {
    sleep($e->getRetryAfter() ?? 30);        // from X-RateLimit-Reset when the server sends it
} catch (QuotaException $e) {
    // usage_limit_exceeded | insufficient_credits — top up, do not retry
} catch (AuthenticationException $e) {
    // invalid_api_key | expired_api_key
} catch (RequestException $e) {
    // your input: invalid_file_type, file_too_large, unknown_model, ...
} catch (ExtractionFailedException $e) {
    // see the billing note below
} catch (JobException $e) {
    // job_expired | job_not_found | chunk_source_lost | chunk_in_progress
} catch (ServerException | TransportException $e) {
    // retryable
}

For finer branching, read the code off the base class:

catch (DocXtractException $e) {
    match ($e->getErrorCode()) {
        'insufficient_credits' => $this->notifyBilling(),
        'unknown_model'        => $this->reportBadModel($e->getDetails()),
        default                => $e->isRetryable() ? $this->requeue() : throw $e,
    };
}
Exception Codes
AuthenticationException invalid_api_key, expired_api_key
QuotaException usage_limit_exceeded, insufficient_credits
RateLimitException rate_limit_exceeded, too_many_open_jobs
RequestException invalid_request, invalid_file, invalid_file_type, file_too_large, invalid_options, unknown_model, page_limit_exceeded, method_not_allowed
ExtractionFailedException extraction_failed
JobException job_not_found, job_expired, chunk_in_progress, chunk_source_lost
ServerException server_error, persist_failed, server_busy
TransportException network failure — the API never answered

An unrecognised code falls back to DocXtractException rather than throwing, so a new server-side code will not break a deployed copy of this SDK.

Billing note. extraction_failed on the synchronous path still deducts 1 credit. Budget for it when reconciling usage. On the multi-page path the chunk stays available and is not charged until it succeeds. persist_failed charges nothing.

Pre-flight validation

File type and size are checked locally before any network call, so an unsupported file costs no round trip and no credits. Accepted: PDF, JPG, PNG, up to 10 MB and 150 pages.

Configuration

$dx = new DocXtract([
    'api_key'        => 'sk_...',                      // required
    'base_url'       => 'https://api.docxtract.io',    // default
    'base_path'      => '/v3.1',                       // default
    'timeout'        => 120,                           // seconds per request
    'max_retries'    => 3,                             // retryable errors only
    'chunk_pause_ms' => 0,                             // delay between chunk calls
    'user_agent'     => '',                            // override the default UA
]);

The base URL has no /api prefix. /api is the server's docroot, not part of the public path. If you see a TransportException about non-JSON output, this is usually why.

Pinning to v3

base_path exists for customers still on the older version:

$dx = new DocXtract(['api_key' => '...', 'base_path' => '/v3']);

v3 has no models and no multi-page support, and accepts PDF only. models() throws a clear error rather than a confusing 404. New integrations should stay on v3.1.

Tests

composer install
composer test

The suite is offline — no API key or network needed. tests/bootstrap.php falls back to a built-in PSR-4 autoloader if vendor/ is absent.

Links

Built by RPATech | docxtract.io