> ## Documentation Index
> Fetch the complete documentation index at: https://daily-docs-flows-declarative.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Eval Suites

> Spawn agents and run many scripted and simulated scenarios concurrently from a single manifest with pipecat eval suite.

`pipecat eval run` tests scenarios against an agent you started yourself. A **suite** goes one step further: you list agents and scenarios in a manifest, and `pipecat eval suite` spawns each agent with its eval transport on its own port, runs its scenarios, tears it down, and aggregates the results, several runs at a time.

Suites are the right tool when you have more than one agent, more than a handful of scenarios, or want a single command for CI. They are also how a [simulation](/pipecat/evals/simulated-scenarios) runs more than once: each run gets a fresh agent. Pipecat's own release evals are a manifest with 100+ example agents plus this command.

## The manifest

```yaml manifest.yaml theme={null}
concurrency: 4 # how many runs execute at once
runs_dir: eval-runs # logs + recordings go to <runs_dir>/<timestamp>/
record: false # record conversation audio (audio-mode scenarios)
scenarios_dir: scenarios # scenario names resolve to <dir>/<name>.yaml

# How to start each agent. {python}, {bot}, and {port} are substituted per run.
spawn: "{python} {bot} -t eval --port {port}"

suite:
  - bot: bots/support-agent.py
    scenarios: [greeting, capital_question, multi_turn]
  - bot: bots/sales-agent.py
    scenarios: [greeting, weather_function_call]
  - bot: bots/booking-agent.py
    scenarios: [greeting, simulated/book_table] # a simulation, listed the same way
  - bot: bots/vision-agent.py
    runner_body: scenarios/vision-body.json # optional --runner-body data
    scenarios: [vision_describe]
```

Paths in the manifest (`bots_dir`, `scenarios_dir`, `runs_dir`, the `bot:` entries) resolve relative to the manifest file, so a manifest is portable: check it into your repo and run it from anywhere. A scenario name can include a subfolder (`simulated/book_table`), and a name ending in `.yaml` is a path relative to the manifest.

Scripted and simulated scenarios are listed the same way, and the file says which it is: `turns:` makes it scripted, `persona:` makes it a simulation. Pipecat's release evals keep the two in `scenarios/scripted/` and `scenarios/simulated/` folders, which is a convention worth copying.

Scenarios are reusable across agents. One `greeting` scenario can cover every agent in the suite.

<Note>
  An optional `runner_body:` points at a JSON file passed to the agent as
  `--runner-body`. It supplies session data the agent would normally receive in
  a `/start` request body (for example, a vision agent's image path).
</Note>

## Running a suite

```bash theme={null}
pipecat eval suite manifest.yaml
```

In a terminal, a live dashboard shows each run's status, a running tally, and total time. When piped (in CI, or driven by a coding assistant), it streams one plain result line per run instead. The command exits `0` only if every run passes.

Useful flags:

```bash theme={null}
pipecat eval suite manifest.yaml -p support       # only bots whose path contains "support"
pipecat eval suite manifest.yaml -s greeting      # only the greeting scenario
pipecat eval suite manifest.yaml -k simulation    # only simulations (or: -k script)
pipecat eval suite manifest.yaml -c 8             # 8 runs at a time
pipecat eval suite manifest.yaml -n nightly       # output to eval-runs/nightly/
pipecat eval suite manifest.yaml -a               # record conversation audio
pipecat eval suite manifest.yaml -d               # save full per-pipeline debug logs
pipecat eval suite manifest.yaml -r 5             # run each pair 5 times
```

Everything except the `suite:` list can live in the manifest or be passed on the command line (the command line wins), so a manifest can be as minimal as a `suite:` list.

## Run output

Each invocation writes to `<runs_dir>/<name>/` (a timestamp when `-n` is omitted):

```
eval-runs/20260610_142200/
  logs/
  results.jsonl                                # one line per run
  logs/
    bots_support-agent.py__greeting.log        # the agent process output
    bots_support-agent.py__greeting.eval.log   # the harness's decision trace
    bots_support-agent.py__greeting.debug.log  # per-pipeline harness logs (-d only)
  recordings/
    bots_support-agent.py__greeting.wav        # conversation audio (record: true or -a)
```

`results.jsonl` carries one line per run, and every line names its `kind`, `script` or `simulation`. A scripted record carries the bot, scenario, attempt, outcome, duration, failures (each with a `kind`), per-turn results, and paths to its artifacts. A simulation record carries the outcome (`passed`, `succeeded`, `error`), how the run ended (`ended_by`), the persona's turn count, every metric's score, value, `failure_kind`, and per-turn verdicts (`yes`, `no`, or `none` when the judge gave none), the judge's reason, the persona's own `end_call` claim, and the whole conversation as `messages`. Lines are appended as each run finishes, so an interrupted sweep keeps everything already done. Runs that didn't pass also carry `events_seen`.

Each run executes in its own process, so a harness that loads local audio models does so on its own, and a crash in one run doesn't stop the others.

<Note>
  `results.jsonl` is written by `pipecat eval suite`. `pipecat eval run` doesn't
  produce one.
</Note>

When a run fails, start with the `.eval.log` decision trace: it's a timestamped record of every event the harness saw, what it matched, what the judge said, and why an assertion failed. The agent's own log sits next to it.

## Testing one agent with many scenarios

If you just want to run a batch of scenarios against an agent you already have running, you don't need a manifest. `pipecat eval run` accepts multiple scenario files and shares the suite's dashboard and tally:

```bash theme={null}
pipecat eval run scenarios/*.yaml --bot-url ws://localhost:7860
pipecat eval run scenarios/ --bot-url ws://localhost:7860       # the whole directory
```

A directory expands to its `.yaml` and `.yml` files in filename order, non-recursively, and files and directories can be mixed in one invocation. A directory holding no scenarios is an error rather than an empty run.

By default the agent is left running afterward so it can serve more evals; pass `--stop-bot` to shut it down when the batch finishes.

## Running a simulation more than once

A persona doesn't say the same thing twice, so one run of a simulation proves little. A simulation's own `runs:` field tells the suite how many times to run it, and every run must pass. The suite reports a pass rate per simulation with a ✓ or ✗ beside it, and exits `1` when any run failed:

```
  Pass rate:
  bots/booking-agent.py  simulated/book_table      3/3 (100%)  ✓  ~12.4s each
```

A run that errored (the agent never came up, the persona's LLM failed, the judge gave no verdict on the goal) is reported but kept out of the rate: it says nothing about whether the agent did its job.

## Repeating a run

A behavior with a race in it passes sometimes. `repeat:` in the manifest, or `--repeat` / `-r` on the command line, runs every (bot, scenario) pair N times and reports a pass rate per pair instead of a single verdict:

```yaml theme={null}
# manifest.yaml
repeat: 5
suite:
  - bot: bots/support-agent.py
    scenarios: [interruption]
```

Attempts interleave across bots rather than running back to back, and artifact filenames gain an attempt suffix (`__001`) only when `repeat` is above 1, so a single pass keeps the filenames it always had.

<Warning>
  A repeated sweep always exits `0`. A pass rate isn't a pass or a fail, so the
  threshold is yours to choose: read `results.jsonl` and decide. Leave `repeat`
  unset for the CI gate below.
</Warning>

A `repeat` set on the command line or in the manifest applies to every run, simulations included, even when it is `1`. It overrides a simulation's `runs:` and turns the requirement into a measurement: rates are reported and the exit code stays `0`. Leave `repeat` unset to let each simulation run its own `runs:` and gate on them.

## Suites in CI

The exit code makes suites CI-ready with no extra glue:

```yaml theme={null}
# e.g. GitHub Actions
- name: Run behavioral evals
  run: pipecat eval suite manifest.yaml
```

For deterministic, key-free CI runs, prefer text-mode scripted scenarios and an OpenAI-compatible judge endpoint you control. Simulations run their persona on the local Ollama model by default, so they need no extra key, but they say something different each run by design, so give them `runs: 3` and expect them to take longer. `-k script` runs the scripted half alone when you want a fast gate on every push and the simulations on a schedule. Audio-mode scenarios work in CI too, but need the harness's TTS and STT services available (local models by default, which also need more CPU).
