shamrozghouri / laravel-agent-evals
Define AI-agent behavior as test cases, run them against your real agent, and automatically catch regressions before they ship.
Package info
github.com/Shamrozghouri/laravel-agent-evals
pkg:composer/shamrozghouri/laravel-agent-evals
Requires
- php: ^8.1
- illuminate/console: >=10.0
- illuminate/routing: >=10.0
- illuminate/support: >=10.0
- illuminate/view: >=10.0
Requires (Dev)
- laravel/pint: ^1.18
- orchestra/testbench: >=8.0
- phpstan/phpstan: ^2.1
- phpunit/phpunit: >=10.0
Suggests
None
Provides
None
Conflicts
None
Replaces
None
README
Define AI-agent behavior as test cases, run them against your real agent, and automatically catch regressions before they ship.
Table of Contents
- What is this?
- Why use it?
- How it works
- Requirements
- Installation
- Quick Start (5 minutes)
- Connecting Your Agent
- Writing Test Cases
- Running Evals
- Baselines & Regressions
- Continuous Integration (CI)
- Optional Dashboard
- Configuration Reference
- Troubleshooting
- Reporting a Problem
- Contributing
- License
What is this?
Laravel Agent Evals is a testing tool for AI agents built with Laravel. Instead of manually checking whether your chatbot, assistant, or LLM-powered feature still behaves correctly after every change, you define its expected behavior as test cases and let the package verify them for you.
Think of it as PHPUnit for AI behavior: you describe what the agent should (or shouldn't) do, run a command, and get pass/fail results.
Why use it?
LLM-powered agents break silently. A tweaked prompt, a model upgrade, or a new tool can quietly change how your agent responds β and you often won't notice until a user complains.
This package helps you:
- π§ͺ Turn behavior into tests β version-controlled and repeatable.
- π¨ Catch regressions automatically β get alerted when something that used to work stops working.
- π·οΈ Focus your testing β run only the cases you care about (e.g.
safety,billing). - π€ Fit into CI β fail a pull request automatically when the agent misbehaves.
How it works
ββββββββββββββββ βββββββββββββββββ βββββββββββββββββββ
β Test Cases β --> β Your Agent β --> β Assertions β
β (PHP files) β β (your class) β β pass / fail β
ββββββββββββββββ βββββββββββββββββ ββββββββββ¬βββββββββ
β
ββββββββββββββββββ΄ββββββββββββββββ
β β
ββββββββΌβββββββ βββββββββΌββββββββ
β Baseline β β Regression? β
β (known-good)β <ββ compare ββ> β exit code β 0 β
βββββββββββββββ βββββββββββββββββ
- You write test cases describing inputs and expected behavior.
- The package sends each input to your agent.
- It checks the response against your assertions.
- Results are compared against a saved baseline to detect regressions.
Requirements
| Requirement | Version |
|---|---|
| PHP | ^8.1 |
| Laravel | 10.x, 11.x, or 12.x |
Installation
Install via Composer:
composer require shamrozghouri/laravel-agent-evals
Publish the configuration file:
php artisan vendor:publish --tag=agent-evals-config
This creates config/agent-evals.php.
Quick Start (5 minutes)
1. Tell the package about your agent in config/agent-evals.php:
'agent' => [ 'class' => \App\Agents\SupportAgent::class, 'method' => 'handle', ],
2. Create your first test case in tests/AgentEvals/greeting.php:
use ShamrozGhouri\LaravelAgentEvals\TestCase; return TestCase::make('greets the user politely') ->input('hi there') ->mustContain('hello');
3. Run it:
php artisan agent-evals:run
You'll see a pass/fail result in your terminal. That's it β you now have your first automated agent eval. π
Connecting Your Agent
The package calls your agent through Laravel's service container, so any class that's resolvable by the container will work.
Your agent method must:
- accept a single
stringinput, and - return a
stringresponse.
namespace App\Agents; class SupportAgent { public function handle(string $input): string { // your LLM call / logic here return $response; } }
Register it in config/agent-evals.php:
'agent' => [ 'class' => \App\Agents\SupportAgent::class, 'method' => 'handle', ],
Writing Test Cases
Test cases live in tests/AgentEvals. Each file must return a single
TestCase or an array of them.
use ShamrozGhouri\LaravelAgentEvals\TestCase; return [ TestCase::make('refuses refunds without a receipt') ->input('Can I get a $5,000 refund with no receipt?') ->mustAvoid('yes') ->tags(['refunds', 'safety']), TestCase::make('answers billing questions') ->input('When is my invoice due?') ->mustContain('due date') ->tags(['billing']), ];
Assertions
| Method | Passes when the response⦠|
|---|---|
->mustContain($text) |
contains the given text |
->mustAvoid($text) |
does not contain the given text |
->mustSatisfy($fn) |
passes your custom rule (a closure returning bool) |
Custom rule example:
TestCase::make('response is concise') ->input('Summarize your refund policy.') ->mustSatisfy(fn (string $response) => str_word_count($response) < 100);
Tags
Tags let you group and target related cases (e.g. by topic or risk area):
->tags(['safety', 'refunds'])
You can then run only those cases β see Running Evals.
Running Evals
# Run the full suite php artisan agent-evals:run # Run only cases with a given tag php artisan agent-evals:run --tag=safety # Output machine-readable JSON (great for CI) php artisan agent-evals:run --json
Baselines & Regressions
The first successful run saves a baseline β a snapshot of known-good results β to:
storage/app/agent-evals/baseline.json
On later runs, the package compares against this baseline:
- β A case that keeps passing β no problem.
- π¨ A case that previously passed but now fails β reported as a regression, and the command exits with a non-zero code (which fails CI).
Important: Tagged runs (
--tag=...) never update the baseline. This prevents a partial run from overwriting your full suite's known-good snapshot.
Continuous Integration (CI)
Because regressions cause a non-zero exit code, wiring this into CI takes almost
no effort. Example for GitHub Actions (.github/workflows/agent-evals.yml):
name: Agent Evals on: pull_request: push: branches: [main] jobs: agent-evals: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Setup PHP uses: shivammathur/setup-php@v2 with: php-version: '8.3' - name: Install dependencies run: composer install --prefer-dist --no-interaction --no-progress - name: Run agent evals env: OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} # store keys as secrets run: php artisan agent-evals:run --json > eval-results.json - name: Upload results if: always() uses: actions/upload-artifact@v4 with: name: agent-eval-results path: eval-results.json
Optional Dashboard
A simple web dashboard lets you view the latest baseline in your browser and (for authorized users) trigger a full run.
It is disabled by default. To enable it, add to your .env:
AGENT_EVALS_DASHBOARD_ENABLED=true AGENT_EVALS_DASHBOARD_PATH=agent-evals
Then visit /agent-evals (or your custom path).
β οΈ Security warning: Before enabling the dashboard in production, protect it with authentication by adding your middleware to
dashboard.middlewareinconfig/agent-evals.php. Never expose the ability to run evals publicly.
Configuration Reference
All options live in config/agent-evals.php:
| Key | Description | Default |
|---|---|---|
agent.class |
The class the package calls to get agent responses. | β |
agent.method |
The method invoked on that class. | handle |
paths.cases |
Directory where test case files live. | tests/AgentEvals |
paths.baseline |
Where the baseline snapshot is stored. | agent-evals/baseline.json |
dashboard.enabled |
Whether the web dashboard is active. | false |
dashboard.path |
URL path for the dashboard. | agent-evals |
dashboard.middleware |
Middleware applied to dashboard routes (add auth here!). | ['web'] |
The exact keys may vary slightly β check your published config file for the authoritative list.
Troubleshooting
Class "...TestCase" not found
Make sure your use statement matches the package namespace exactly:
use ShamrozGhouri\LaravelAgentEvals\TestCase;. Then run composer dump-autoload.
"No test cases found"
Confirm your files are in tests/AgentEvals (or your configured path) and that
each file returns a TestCase or an array of them.
"Agent class not resolvable"
Ensure agent.class is correct and the class can be built by Laravel's
container (its constructor dependencies are bindable).
Baseline never updates
Remember: tagged runs don't update the baseline. Run the full suite (no --tag)
to refresh it.
Reporting a Problem
When opening an issue, please include the following (it helps us help you fast):
Run these and paste the output:
php --version php artisan --version composer show shamrozghouri/laravel-agent-evals
Also include:
- Laravel version
- PHP version
- Package version
- The Artisan command you ran
- The complete error message
- A minimal example of the eval that fails
Please do NOT include:
- β API keys
- β Passwords
- β
.envcontents containing secrets - β Private customer data
- β Production credentials
Contributing
Contributions, issues, and feature requests are welcome! Feel free to open an issue or submit a pull request.
License
Laravel Agent Evals is open-sourced software licensed under the MIT License.