Search by

mahdisphp / laravel-hack-auditor

mahdisphp

AI-assisted security auditor and CTF generator for Laravel, with a deterministic, zero-cost access-control engine. Every claim backed by evidence.

Package info

github.com/mahdi-salmanzade/laravel-hack-auditor

pkg:composer/mahdisphp/laravel-hack-auditor

Statistics

Installs: 115

Dependents: 0

Suggesters: 0

Stars: 5

Open Issues: 0

v2.3.0 2026-09-28 06:37 UTC

This package is auto-updated.

Last update: 2026-09-28 07:03:58 UTC


README

Laravel Hack Auditor

AI-assisted security audits for Laravel — every claim backed by evidence.

Latest Version Total Downloads Stars

See what a scan report looks like on a deliberately vulnerable controller — no setup, no API key. The demo replays pre-recorded findings (no AI call is made); hack:scan runs the real analysis on your code.

hack:demo in action

composer require mahdisphp/laravel-hack-auditor
php artisan hack:demo

That's it. Two commands: 11 confirmed vulnerabilities and 1 review question on a planted controller, scored with the same formula as a real scan. Pass --copy to copy the share text to your clipboard (nothing is copied without it).

The commands

php artisan hack:demo                   # See it in action (no API key)
php artisan hack:scan                   # Scan YOUR app with AI
php artisan hack:scan --diff --html     # Scan only changed files, export HTML report
php artisan hack:scan --format=sarif    # SARIF 2.1.0 for GitHub code scanning
php artisan hack:ctf sql_injection      # Turn vulns into CTF challenges
php artisan hack:report --latest        # HTML (or --format=sarif|markdown) report from saved scan
php artisan hack:benchmark              # Measure recall (precision/recall/F1) on the labeled corpus
php artisan hack:help                   # Full command reference
php artisan hack:usage                  # Token usage & cost stats
php artisan mcp:start hack-auditor      # Expose the scanner to AI agents (Claude Code, Cursor)

hack:scan finds what PHPStan and Snyk can't:

  • "This endpoint fetches a user by ID but never checks ownership" (IDOR)
  • "Admin check reads is_admin from the request, not the session" (Auth bypass)
  • "Login route has no throttle middleware" (Brute-forceable)
  • "Any authenticated user can set their own plan to 'pro' without payment" (Auth bypass)

20 vulnerability types. OWASP Top 10 mapped, each with a CWE id. Every finding has file, line and an explanation — and a suggested fix only when the scanner can prove every identifier it would name. See Precision.

Deterministic detection engine — not just the AI

Alongside the AI pass, framework-aware detectors run on every scan and merge into the report, giving reproducible coverage that doesn't drift with AI run-to-run variance: IDOR / broken access control (policy-vs-route mismatch, is_admin in $fillable, unauthorized find()/findOrFail() exposure), SSRF (Http::get()/cURL with a user-controlled URL), and sensitive-data exposure (password/token/secret fields returned in a response). These are the OWASP-#1 access-control bugs generic SAST and generic AI both miss because they don't understand Laravel.

The engine parses your code into a real AST (nikic/php-parser) and resolves Laravel semantics before deciding anything: what a receiver's type is, so an Eloquent ->get() and a local service's ->delete() are not mistaken for HTTP calls; which abilities a Policy actually declares, so a missing store ability is not reported as a bypass; and that $request->user() is the authenticated user rather than attacker-controlled input. Earlier versions pattern-matched raw source and got all three wrong.

Measured accuracy — and what the measurement is worth

hack:benchmark runs the scanner against the labeled corpus in tests/Fixtures/benchmark/ and reports precision / recall / F1 overall and per type, usable as a CI gate (--min-f1). It ships in the repo — run it yourself.

Read the number for what it is. That corpus is synthetic: every sample was authored alongside the detectors, and the routes that expose them are declared in the corpus's own routes.php. It is a recall check and a regression gate — "does the engine still find the bugs it is supposed to find" — not a precision claim about real code. Nothing in it has the ambiguity, framework idiom or deliberate-by-design pattern that produces false positives in a real application. The command prints this caveat with every result so a number copied out of a terminal carries its own scope. Real-code precision is measured separately, against unmodified third-party Laravel applications.

php artisan hack:benchmark --deterministic   # reproducible engine only — no AI key, no network
php artisan hack:benchmark --min-f1=0.9      # full pipeline (AI + deterministic), CI gate

--deterministic scores only the labels the provider-independent detectors own, so the reproducible half of the scanner can be gated on every commit with no API key.

Why the corpus ships a route manifest. The access-control engine refuses to report a record exposure it cannot attribute to a routed entry point — the rule that took 191 false IDOR reports on 6,221 real files down to zero. Standalone fixture files have no application around them, so without a route table every sample resolved to unreachable and was dropped before analysis: the gate kept printing a score for detectors it had stopped exercising. tests/Fixtures/benchmark/routes.php gives the corpus what a real app has. The reachability rule is untouched; a controller sample with no route entry now fails the run loudly rather than being silently unmeasured.

Call it from your AI editor

mcp:start hack-auditor exposes the scanner as MCP tools (scan_path, scan_diff, explain_finding) so Claude Code, Cursor, and other agents can run a real taint-aware Laravel audit mid-edit instead of guessing.

Precision, measured on real code

The measurement. The deterministic engine was run over 6,221 files from six large open-source Laravel applications — Monica, Akaunting, Pixelfed, BookStack, Snipe-IT, Koel — plus the Laravel framework's own src/. None of them was consulted while writing the detectors.

asserted vulnerabilities review items
without a route map 0 28
with a route map (what a real scan has) 0 3

Zero asserted findings on well-maintained real code. Recall was held while getting there: the deliberately-vulnerable laravel-vuln-lab still yields all 7 of its planted access-control findings.

Two finding classes, because certainty and severity are different questions. Severity answers "how bad if real"; it cannot express "how sure am I". Conflating them is how a scanner ships confident nonsense.

  • Confirmed vulnerability — every link of the evidence chain was resolved from your code: an attacker-controlled source, the sink it reaches, and the absence of a guard on the path. Only these are counted, scored, and allowed to fail a build.
  • Needs review — security-sensitive code the scanner cannot prove either way, phrased as a question and excluded from the count, the score and the exit code. A review item never carries a suggested fix.

That second rule is structural, not a convention: a review finding's fix string is dropped in Vulnerability's constructor, so a detector cannot reintroduce one by forgetting. It exists because this tool has shipped fixes that break applications — advising an ability that a policy never declared, and advising the removal of a $fillable column that a multi-tenant app needed to stay tenant-scoped.

A suggested fix must survive every one of these, or none is emitted and the finding explains why instead:

  • the finding is a confirmed vulnerability, never a review item;
  • every identifier it names was resolved from the analysed file — policy classes are quoted from the class actually resolved, never synthesised as {Model}Policy;
  • every variable it names is definitely assigned on all paths to the insertion point, so a binding inside a try/catch, a match arm or a loop body is never named;
  • the method it advises calling is actually callable on that class.

What we cannot see. Route middleware registered at runtime, dynamically resolved policies, gates defined in service providers, anything reached via __call, and authorization enforced outside the analysed file set. That is exactly why the review class exists.

Multi-pass verification (v1.6)

Pass --verify to have the AI attempt a concrete exploit for every HIGH or CRITICAL finding. Findings the model can exploit retain their severity and ship with a copy-paste exploit payload (exploit_proof). Findings it cannot exploit are downgraded one tier (Critical→High, High→Medium) with the original severity preserved in original_severity for audit trail — a placeholder or hedging response is treated as no-exploit.

php artisan hack:scan --verify
# → Verification 8/8 HIGH+ findings had working exploits (0 downgraded)
#   Verification tokens: 15,288 input + 1,702 output = 16,990 total

⚠️ --verify approximately doubles API cost on scans with many HIGH+ findings. Recommended for pre-release audits, not every CI run. Enable by default via HACK_AUDITOR_VERIFY=true.

Technical failures (AI timeouts, malformed responses) leave the finding untouched rather than downgrading on noise. The JSON output gains a verification sub-object with verified/downgraded counts and a separate token bucket so pass-1 and pass-2 cost are distinguishable.

Token usage & cost tracking

Every scan shows token consumption and estimated cost. Auto-detects your AI provider's pricing from a built-in registry of 30+ models (Anthropic, OpenAI, Gemini, xAI, Ollama). Budget your scans with --limit.

Token Usage ...... 97,188 prompt + 3,080 completion = 100,268 total
AI Requests ...... 7
Estimated Cost ... $0.5629
Model ............ claude-opus-5 (anthropic)

Quick setup (2 minutes)

php artisan install:ai                  # Install Laravel AI

Add one API key to .env:

ANTHROPIC_API_KEY=sk-ant-your-key-here  # or OPENAI_API_KEY, or GEMINI_API_KEY

Scan:

php artisan hack:scan

Done. The package uses whatever provider you configured in Laravel AI. Optionally override just for this package:

HACK_AUDITOR_AI_PROVIDER=anthropic
HACK_AUDITOR_AI_MODEL=claude-opus-5

On sampling parameters. Scans run at a low fixed temperature for reproducibility, but Anthropic removed temperature from Claude Opus 4.7 onward — sending it to Opus 4.7/4.8, Opus 5, Sonnet 5 or Fable 5 returns HTTP 400. The scanner detects those models and omits the parameter, so they work out of the box; on those models reproducibility comes from the deterministic detection engine rather than from temperature.

All scan flags
Flag What it does
--path=app/Http/Controllers Scan a specific directory (walks it recursively) or a single file. A missing or refused path exits 2
--severity=High Show and gate on High+ only (default: severity.minimum_report, else Low)
--fail-on=high Exit 1 when a confirmed finding at or above this severity remains after --severity and the baseline: critical (default), high, medium, low or none
--format=sarif Output format: table (default), json, sarif (SARIF 2.1.0) or markdown
--json JSON output for CI/CD (alias for --format=json)
--fix Show suggested fixes — only for confirmed findings that carry one
--verify Second AI pass that tries to build an exploit for each HIGH+ finding; unexploitable ones are downgraded a tier
--deterministic Run only the reproducible, provider-independent detectors — no AI key, no network
--html Generate HTML report
--save Save results to JSON file (for hack:report, hack:ctf --from-scan and the "since last scan" delta)
--force Skip confirmation prompt
--detailed Full descriptions, confidence, CWE and taint trace per finding
--diff Only scan git-changed files (great for CI), in one chunked pass. Combine with --path to restrict to a subtree. Works when the app lives in a monorepo subdirectory. A base that cannot be resolved — typically a shallow CI checkout; use fetch-depth: 0 — exits 2 instead of reporting no changes
--base=develop Base branch for --diff
--limit=50000 Cap token budget for the scan
--baseline Require the baseline: exit 2 if the file is missing (it is applied automatically whenever it exists)
--update-baseline Save current findings as baseline
--no-baseline Ignore baseline file

Exit codes: 0 no confirmed finding at or above --fail-on; 1 at least one remains after filters; 2 invalid usage, or the scan target was missing/refused (nothing was scanned). Review items never affect the exit code.

Score: max(0, 100 − Σ severity weight) over confirmed vulnerabilities (critical 40, high 20, medium 10, low 5). It is withheld (null, with a reason) when coverage was incomplete, nothing was analysed, or the scan target could not be resolved.

Report flags
Flag What it does
--latest Generate report from the most recent saved scan (the default)
--id=ULID Generate report from a specific scan ID
--format=html html (default), sarif or markdown
--output=path Custom output file path
Usage flags
Flag What it does
--days=30 Show usage from the last N days (default: 30)
--json Output as JSON
--clear Clear the usage log

Generate CTF challenges from real vulns

Train your team by turning actual findings into Capture The Flag exercises:

php artisan hack:ctf sql_injection    # By type
php artisan hack:ctf --from-scan      # From latest scan results
php artisan hack:ctf --all            # Every finding in the latest saved scan (implies --from-scan)

Each challenge outputs a ready-to-run directory: README, vulnerable code, solution, flag file, and docker-compose.

HTML reports, git-aware scanning, baselines

php artisan hack:scan --html            # Beautiful dark-themed HTML report
php artisan hack:scan --diff            # Only scan files changed in your branch
php artisan hack:scan --update-baseline # Accept current findings as known
php artisan hack:report --latest        # Regenerate report from saved scan

The HTML report is a single self-contained file — dark theme, animated score ring, collapsible cards, copy-paste code blocks, token usage breakdown. Professional enough to attach to a security audit.

--diff scans only what your PR touches. --update-baseline lets teams acknowledge known risks so CI doesn't fail on accepted findings. Baseline entries are matched by a fingerprint — type + file + the normalised content of the flagged line — so an accepted finding stays accepted when the AI rewords it or code above it moves, and two identical flaws in one file stay two findings. Baselines written by older versions are still honoured.

Every finding in JSON, SARIF and MCP output carries that fingerprint, its confidence, its class (vulnerability or review) and references (the CWE entry plus the OWASP Cheat Sheet for its type). JSON output also carries score_breakdown: the per-severity count × weight that produced the score, or null when the score is withheld.

Use it in code

use Mahdi\HackAuditor\Facades\HackAuditor;
use Mahdi\HackAuditor\Support\UsageTracker;

$report = HackAuditor::scan();

if ($report->hasCritical()) {
    // Block deployment, alert Slack, panic, etc.
}

echo $report->overallScore;                  // 0-100
echo $report->criticalCount();               // int
echo $report->getUsageTracker()?->totalTokens();  // tokens used
echo $report->getUsageTracker()?->estimateCost();  // estimated $$$

// Budget-capped scan
$tracker = new UsageTracker(tokenLimit: 50_000);
$report = HackAuditor::scan(tracker: $tracker);

// Scan history
$history = HackAuditor::history();
$latest = $history->latest();                // array or null
$all = $history->recent(10);                 // last 10 scans
CI/CD pipeline example
# .github/workflows/security.yml
name: Security Audit
on: [push, pull_request]
jobs:
  hack-audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: shivammathur/setup-php@v2
        with:
          php-version: '8.3'
      - run: composer install --no-interaction
      - run: php artisan hack:scan --format=sarif --fail-on=high --force > hack-auditor.sarif
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
      - uses: github/codeql-action/upload-sarif@v3
        if: always()
        with:
          sarif_file: hack-auditor.sarif
All configuration options
php artisan vendor:publish --tag=hack-auditor-config
Option Default Description
ai.provider null AI provider override
ai.model null Model override
ai.temperature 0.3 Lower = more deterministic
ai.max_tokens 4096 Max tokens per AI response
ai.timeout 120 HTTP timeout in seconds
scan.paths Controllers, Models, Requests, Middleware, routes What to scan
scan.exclude */vendor/*, */node_modules/*, */tests/* Excluded paths
scan.file_extensions ['.php'] File extensions to scan
scan.max_file_size_kb 500 Skip files larger than this
scan.chunk_size 10 Files per AI request
scan.confirm_above_files 20 Prompt before large scans
scan.sensitive_patterns .env*, *.key, *.pem, storage/logs/* Always excluded
scan.diff_base_branch null Base branch for --diff (auto-detects main/master)
scan.baseline_path base_path('hack-auditor-baseline.json') Path to baseline JSON file
context.enabled true Context-aware scanning (routes, middleware, policies, models)
context.max_context_tokens 8000 Token budget for context
context.include_routes true Include route info in context
context.include_middleware true Include middleware info in context
context.include_rate_limiters true Include registered rate limiters in context
context.include_policies true Include policy info in context
context.include_form_requests true Include form request info in context
context.include_models true Include model info in context
context.include_config true Include security-relevant config values in context
context.include_environment true Include environment facts (e.g. debug mode) in context
context.extra_context_paths [] Additional paths to include in context
severity.minimum_report 'Low' Default for hack:scan --severity (also scopes the --fail-on gate)
ctf.output_path hack-auditor/ctf CTF output directory
report.output_path hack-auditor/reports HTML report output directory
share.default_hashtags ['#LaravelSecurity', '#HackAuditor', '#CTF'] Hashtags for sharing
share.ai_tweets true AI-generated share text
usage.default_limit 0 Default --limit value (0 = unlimited)
usage.cost_per_1m_input 3.00 Cost per 1M input tokens
usage.cost_per_1m_output 15.00 Cost per 1M output tokens
usage.show_usage true Show token usage after scan
usage.log_enabled true Auto-log usage to storage/hack-auditor/usage.json
verification.enabled false Run --verify on every scan (HACK_AUDITOR_VERIFY)

Zero database dependencies. All data stored as JSON files in storage/hack-auditor/.

Security

This package sends source code to AI providers. Files matching .env*, *.key, *.pem, and storage/logs/* are always excluded. Review your provider's data retention policies.

Found a vulnerability in this package? Email mahdi@mindzone.tech.

Contributing

PRs welcome. Run composer test and vendor/bin/pint before submitting.

License

MIT — LICENSE

If this saved you from getting hacked, star the repo.