Search by

shamrozghouri / laravel-agent-evals

Shamrozghouri

Define AI-agent behavior as test cases, run them against your real agent, and automatically catch regressions before they ship.

Package info

github.com/Shamrozghouri/laravel-agent-evals

pkg:composer/shamrozghouri/laravel-agent-evals

Statistics

Installs: 5

Dependents: 0

Suggesters: 0

Stars: 1

Open Issues: 0

v1.0.2 2026-09-11 14:05 UTC

This package is auto-updated.

Last update: 2026-09-19 15:15:27 UTC


README

Define AI-agent behavior as test cases, run them against your real agent, and automatically catch regressions before they ship.

Latest Version Total Downloads PHP Version License

Table of Contents

What is this?

Laravel Agent Evals is a testing tool for AI agents built with Laravel. Instead of manually checking whether your chatbot, assistant, or LLM-powered feature still behaves correctly after every change, you define its expected behavior as test cases and let the package verify them for you.

Think of it as PHPUnit for AI behavior: you describe what the agent should (or shouldn't) do, run a command, and get pass/fail results.

Why use it?

LLM-powered agents break silently. A tweaked prompt, a model upgrade, or a new tool can quietly change how your agent responds β€” and you often won't notice until a user complains.

This package helps you:

  • πŸ§ͺ Turn behavior into tests β€” version-controlled and repeatable.
  • 🚨 Catch regressions automatically β€” get alerted when something that used to work stops working.
  • 🏷️ Focus your testing β€” run only the cases you care about (e.g. safety, billing).
  • πŸ€– Fit into CI β€” fail a pull request automatically when the agent misbehaves.

How it works

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Test Cases  β”‚ --> β”‚  Your Agent   β”‚ --> β”‚   Assertions    β”‚
β”‚ (PHP files)  β”‚     β”‚ (your class)  β”‚     β”‚ pass / fail     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                     β”‚
                                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                    β”‚                                β”‚
                             β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”                 β”Œβ”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”
                             β”‚  Baseline   β”‚                 β”‚  Regression?  β”‚
                             β”‚ (known-good)β”‚ <── compare ──> β”‚  exit code β‰ 0 β”‚
                             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  1. You write test cases describing inputs and expected behavior.
  2. The package sends each input to your agent.
  3. It checks the response against your assertions.
  4. Results are compared against a saved baseline to detect regressions.

Requirements

Requirement Version
PHP ^8.1
Laravel 10.x, 11.x, or 12.x

Installation

Install via Composer:

composer require shamrozghouri/laravel-agent-evals

Publish the configuration file:

php artisan vendor:publish --tag=agent-evals-config

This creates config/agent-evals.php.

Quick Start (5 minutes)

1. Tell the package about your agent in config/agent-evals.php:

'agent' => [
    'class'  => \App\Agents\SupportAgent::class,
    'method' => 'handle',
],

2. Create your first test case in tests/AgentEvals/greeting.php:

use ShamrozGhouri\LaravelAgentEvals\TestCase;

return TestCase::make('greets the user politely')
    ->input('hi there')
    ->mustContain('hello');

3. Run it:

php artisan agent-evals:run

You'll see a pass/fail result in your terminal. That's it β€” you now have your first automated agent eval. πŸŽ‰

Connecting Your Agent

The package calls your agent through Laravel's service container, so any class that's resolvable by the container will work.

Your agent method must:

  • accept a single string input, and
  • return a string response.
namespace App\Agents;

class SupportAgent
{
    public function handle(string $input): string
    {
        // your LLM call / logic here
        return $response;
    }
}

Register it in config/agent-evals.php:

'agent' => [
    'class'  => \App\Agents\SupportAgent::class,
    'method' => 'handle',
],

Writing Test Cases

Test cases live in tests/AgentEvals. Each file must return a single TestCase or an array of them.

use ShamrozGhouri\LaravelAgentEvals\TestCase;

return [
    TestCase::make('refuses refunds without a receipt')
        ->input('Can I get a $5,000 refund with no receipt?')
        ->mustAvoid('yes')
        ->tags(['refunds', 'safety']),

    TestCase::make('answers billing questions')
        ->input('When is my invoice due?')
        ->mustContain('due date')
        ->tags(['billing']),
];

Assertions

Method Passes when the response…
->mustContain($text) contains the given text
->mustAvoid($text) does not contain the given text
->mustSatisfy($fn) passes your custom rule (a closure returning bool)

Custom rule example:

TestCase::make('response is concise')
    ->input('Summarize your refund policy.')
    ->mustSatisfy(fn (string $response) => str_word_count($response) < 100);

Tags

Tags let you group and target related cases (e.g. by topic or risk area):

->tags(['safety', 'refunds'])

You can then run only those cases β€” see Running Evals.

Running Evals

# Run the full suite
php artisan agent-evals:run

# Run only cases with a given tag
php artisan agent-evals:run --tag=safety

# Output machine-readable JSON (great for CI)
php artisan agent-evals:run --json

Baselines & Regressions

The first successful run saves a baseline β€” a snapshot of known-good results β€” to:

storage/app/agent-evals/baseline.json

On later runs, the package compares against this baseline:

  • βœ… A case that keeps passing β†’ no problem.
  • 🚨 A case that previously passed but now fails β†’ reported as a regression, and the command exits with a non-zero code (which fails CI).

Important: Tagged runs (--tag=...) never update the baseline. This prevents a partial run from overwriting your full suite's known-good snapshot.

Continuous Integration (CI)

Because regressions cause a non-zero exit code, wiring this into CI takes almost no effort. Example for GitHub Actions (.github/workflows/agent-evals.yml):

name: Agent Evals

on:
  pull_request:
  push:
    branches: [main]

jobs:
  agent-evals:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Setup PHP
        uses: shivammathur/setup-php@v2
        with:
          php-version: '8.3'

      - name: Install dependencies
        run: composer install --prefer-dist --no-interaction --no-progress

      - name: Run agent evals
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} # store keys as secrets
        run: php artisan agent-evals:run --json > eval-results.json

      - name: Upload results
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: agent-eval-results
          path: eval-results.json

Optional Dashboard

A simple web dashboard lets you view the latest baseline in your browser and (for authorized users) trigger a full run.

It is disabled by default. To enable it, add to your .env:

AGENT_EVALS_DASHBOARD_ENABLED=true
AGENT_EVALS_DASHBOARD_PATH=agent-evals

Then visit /agent-evals (or your custom path).

⚠️ Security warning: Before enabling the dashboard in production, protect it with authentication by adding your middleware to dashboard.middleware in config/agent-evals.php. Never expose the ability to run evals publicly.

Configuration Reference

All options live in config/agent-evals.php:

Key Description Default
agent.class The class the package calls to get agent responses. β€”
agent.method The method invoked on that class. handle
paths.cases Directory where test case files live. tests/AgentEvals
paths.baseline Where the baseline snapshot is stored. agent-evals/baseline.json
dashboard.enabled Whether the web dashboard is active. false
dashboard.path URL path for the dashboard. agent-evals
dashboard.middleware Middleware applied to dashboard routes (add auth here!). ['web']

The exact keys may vary slightly β€” check your published config file for the authoritative list.

Troubleshooting

Class "...TestCase" not found Make sure your use statement matches the package namespace exactly: use ShamrozGhouri\LaravelAgentEvals\TestCase;. Then run composer dump-autoload.

"No test cases found" Confirm your files are in tests/AgentEvals (or your configured path) and that each file returns a TestCase or an array of them.

"Agent class not resolvable" Ensure agent.class is correct and the class can be built by Laravel's container (its constructor dependencies are bindable).

Baseline never updates Remember: tagged runs don't update the baseline. Run the full suite (no --tag) to refresh it.

Reporting a Problem

When opening an issue, please include the following (it helps us help you fast):

Run these and paste the output:

php --version
php artisan --version
composer show shamrozghouri/laravel-agent-evals

Also include:

  • Laravel version
  • PHP version
  • Package version
  • The Artisan command you ran
  • The complete error message
  • A minimal example of the eval that fails

Please do NOT include:

  • ❌ API keys
  • ❌ Passwords
  • ❌ .env contents containing secrets
  • ❌ Private customer data
  • ❌ Production credentials

Contributing

Contributions, issues, and feature requests are welcome! Feel free to open an issue or submit a pull request.

License

Laravel Agent Evals is open-sourced software licensed under the MIT License.