Skip to main content

Command Palette

Search for a command to run...

How to Set Up Evals for a Complex Multi-Agent System

Using Codex and Promptfoo to define success, diagnose failures, and choose models

Updated
•25 min read•View as Markdown
How to Set Up Evals for a Complex Multi-Agent System
G
Applied AI Engineer at OpenAI. I write about building AI applications, evaluating models, and turning experiments into practical solutions, drawing on my background in cloud architecture and software engineering.

Evaluating a multi-agent application starts with a practical question: did the system do the job? Answering it takes more than checking the final response. A coordinator might follow a specialist's incorrect advice. A specialist might produce a sound recommendation from incomplete evidence. Several individually capable agents might combine into a workflow that repeats work, misses a handoff, or declares success too early.

A useful evaluation makes those failures visible and helps you decide what to change. That might be a model or reasoning setting. It might also be a tool, an instruction, or the application's definition of completion.

The evaluation unit here is an agentic episode: a task unfolding through decisions, tool calls, returned evidence, and a final answer or change in system state. One test can exercise several dependent steps. The graders inspect properties of that episode: whether the agent retrieved the right evidence, respected its authority and budget, made supported decisions, and reached the required outcome. A simple assertion can detect a consequential failure within that multi-step behavior.

Animated evaluation workflow with a feedback loop: diagnose application defects, harden the app, add regression tests, and recheck before relying on model choices.

Figure 1. Start with realistic examples and reviewed expectations. Pilot and comparison results can expose application defects: diagnose them, harden the application, and add regression tests before validating again. Keep model choices provisional until the assembled workflow meets its requirements.

Where Codex and Promptfoo fit

This guide describes how I would build that evaluation with Codex and Promptfoo. Promptfoo is an open-source toolkit for evaluating AI applications and testing them for security weaknesses. You define test cases, choose the models or application endpoints to exercise, and specify checks called assertions. Promptfoo runs those tests, records outputs and scores, and provides a web viewer for comparing results side by side. Its command-line tool and library let those evaluations become part of a development workflow.

Promptfoo charts and five model/reasoning columns comparing saved Risk Review results on nine cases.

Promptfoo compares configurations on the same cases and lets you inspect the outputs and grades behind each result.

Promptfoo is now part of OpenAI, following the acquisition announced in March 2026, and remains open source. This article focuses on its evaluation capabilities: using a custom adapter to exercise tool-using agents, grade their episodes, and inspect the results.

Codex helps inspect the application, build the recording and replay code, and investigate failures; Promptfoo provides the test runner, assertion framework, and comparison interface. The engineering team still needs to define success, establish evidence of correctness, and decide how the results will support a configuration choice.

The running example: Mars Mission Control

The running example is a five-agent mission simulation evaluated with Promptfoo and built with the OpenAI Agents API. The workflow below applies to other systems that delegate work, retrieve evidence, and coordinate actions.

In this application, the Mission Director selects specialists, combines their evidence, and proposes a response for human authorization. Life Support assesses cabin conditions and air processing. Power & Thermal evaluates electrical reserves and cooling. Weather & Navigation assesses hazards, navigation, and crew location. Risk Review challenges assumptions, identifies evidence gaps, and flags unsafe trade-offs. These different responsibilities are why each role needs its own cases and grading criteria: a correct resource calculation, a well-supported challenge, and a coordinated plan are different contributions to the same outcome.

The agents work with several kinds of evidence: timestamped telemetry and available recent readings; safety protocols specifying permitted actions, durations, and prerequisites; measured objectives and command execution results; and simulated orbital, maintenance, and crew reports. These reports cross-check the same simulator observations rather than provide independent sensor measurements. Specialists can also use an arithmetic tool to check numerical forecasts, while the Director combines their reports into a proposed plan.

The complexity lies in connecting that evidence under constraints. A proposed crew return may depend on mobility, communications, remaining power, and the hazard deadline. A calculation can be numerically correct but rely on the wrong assumptions or omit a prerequisite. The eval therefore needs to check evidence retrieval, calculations, handoffs, and authority boundaries, in addition to the recommendation. Similar dependencies arise when business agents combine customer records, policy documents, inventory, and tool results to decide what can actually be done.

Ares-7 mission dashboard showing an approaching dust storm, the response window, and habitat and crew telemetry.

The mission app brings together incident conditions, telemetry, and time pressure for its five agents to assess.

Saved mission assessment showing specialist recommendations returning to the Mission Director beside the agent interaction map.

In a saved assessment, specialists return evidence-backed recommendations to the Mission Director, who submits a coordinated plan for human review.

1. Map the workflow and define success

Begin with one representative user task. Follow it from the initial request through the coordinator, specialists, tools, and final outcome. For each agent, record its responsibility, inputs, available tools, expected output, and the agent or component that consumes that output.

Multi-agent workflow with component colors and numbered boundaries for one agent, one handoff, and the entire workflow, including authorization, execution, verification, and failure handling.

The system under test. Colors identify component types; numbered boundaries identify evaluation scope. The outer boundary encompasses all components and branches, including authorization, execution, verification, and failure handling. Smaller tests isolate an agent's episode or a handoff; the coordinator can also be tested as an individual agent. Exercise different paths across separate cases, using controlled approve/decline inputs to test authorization behavior.

Then write the success criteria at two levels: what each agent must contribute, and what the complete system must accomplish. These can differ substantially. A research agent may need to return supported claims and citations; the system may need to produce a decision that resolves conflicting evidence. A support specialist may identify the right action; the system still needs to execute the authorized action and verify completion.

Specify important constraints alongside the outcome: permissions, deadlines, tool budgets, required evidence, and acceptable behavior when the request cannot be completed. Include at least one plausible response that should fail. This often exposes ambiguity more quickly than another example of a good answer.

Ask Codex to inspect the real entry points and tool implementations before proposing tests. Check the documentation against the execution path so the tests exercise the intended behavior.

The first deliverable is a workflow map and a written success contract. Have the person responsible for the product's behavior review that contract before encoding it in a grader.

2. Evaluate agents, handoffs, and the complete workflow

Use three complementary test scopes:

Scope What you exercise What it helps diagnose
Individual agent One role, with controlled input and its permitted tools Reasoning, retrieval, calculations, and role compliance
Handoff A producer's output as the next agent's input Missing information, incompatible structure, lost uncertainty, and conflicting advice
Complete workflow The actual agents working together from a user request Routing, coordination, recovery, resource use, and the final outcome
Three evaluation scopes: individual agents, handoffs between agents, and complete workflows.

Figure 2. Each scope answers a different question. Component tests help locate defects; workflow tests establish whether the parts work together.

An individual-agent test can still contain multiple steps. The agent might retrieve a record, call a calculator, and then respond. Treat that bounded sequence as one episode, and retain its tool activity as well as its answer. Our model-selection comparisons used the following scope: each case involved one agent, its tools, and its final response. A separate end-to-end follow-up, described below, exercised the final selected team on new simulated conditions.

For a fair role comparison, give competing configurations the same starting input. A coordinator can receive a frozen set of specialist reports so its ability to use those reports is measured consistently. Include imperfect or contradictory reports when handling them is part of its job. In our replay, consultation tools returned recorded reports; they did not launch fresh specialists or answer new hypothetical questions. That boundary limits what the component test establishes.

Separately, test the selected agents together. Replayed handoffs isolate a component; newly generated handoffs expose interactions. A collection of passing component tests does not establish that the complete workflow succeeds.

In Promptfoo, give each role its own review view and retain a consolidated overview for navigation. A reader should be able to find the relevant cases without sorting through every experiment. Keep phase labels when the tools or grading contract changed.

Promptfoo evaluation library showing five per-agent views and a consolidated overview.

Promptfoo in practice. Five role views each contain nine cases; the overview brings their saved results together. This is an index of staged experiments, not a new combined-team test. Aggregate pass rates are hidden here because configurations cover different roles and phases.

3. Record realistic interactions and build reviewed cases

Instrument the application before collecting a large dataset. Record at the boundary where requests and tool calls occur so you retain what the agent received rather than reconstructing it later.

For each episode, capture the task, role instructions, model and reasoning setting, tool definitions, tool calls and returned evidence, final output, errors, elapsed time, and available usage. Preserve the initial application state needed for replay. Redact sensitive information and keep credentials out of fixtures.

Use real traces to find representative tasks and difficult interactions. If the project is new, begin with explicitly authored scenarios based on requirements, then add real interactions as the application becomes usable. Include ordinary work, boundary conditions, missing information, infeasible requests, and tool failures. Several repetitions of an easy task do not substitute for this variety.

Turn each selected example into a case with three parts:

  • Input: the task, starting context, state, and tools available to the agent.

  • Expected behavior: required outcomes, acceptable alternatives, evidence requirements, and material mistakes.

  • Provenance: where the case came from and which application, tool, and rubric versions it uses.

For example, an illustrative support-agent case could start with a customer requesting a refund, a verified order record, and an expired eligibility window. The expected behavior is to explain the applicable policy and use the allowed escalation path. Issuing an unauthorized refund is a critical failure. An agent test checks that decision; a handoff test checks that the coordinator preserves the eligibility constraint; a workflow test checks that the promised escalation actually exists in the resulting system state.

A recorded answer can be wrong. A “golden” dataset needs reviewed expectations grounded in business rules, authoritative sources, checked calculations, or expert decisions. A stronger model can draft reference answers and suggest overlooked cases; its confidence does not certify the labels. Our reference and failure reviews were Codex-assisted development reviews, not independent expert adjudication.

For open-ended work, describe acceptable behavior rather than requiring one exact reference response. A valid alternative plan should pass when it satisfies the same requirements.

Keep some representative cases out of development. Once a case guides a prompt, tool, or rubric repair, it belongs to the development set; it no longer provides untouched validation evidence.

4. Connect the application to Promptfoo

Build an adapter, called a custom provider in Promptfoo, that invokes the application's agent execution code. Declare any evaluation-specific substitutions, such as fixed specialist reports or controlled tool responses. The adapter connects a test case to an agent episode.

It should load the fixture, create an isolated session or reset the state, apply the requested model and reasoning settings, execute the agent with bounded time and tool use, and save the result along with its trace. Return enough information for assertions to inspect both the answer and the workflow. A timeout or tool failure should still be recorded as an attempt. Specify whether tools use frozen responses, a simulator, or live test services. Isolate writes from production, and reset mutable state between attempts; otherwise one candidate can change the conditions for the next.

For a multi-agent application, organize the harness around the test scopes above. Role tests supply controlled upstream inputs. Handoff tests exercise the producer–consumer boundary. Workflow tests run the actual orchestration and check the resulting state.

Use a dedicated evaluation directory to keep cases, reviewed expectations, the adapter, assertions, configuration, saved outputs, and reports together. Version the test contract with the application. Check that the requested reasoning setting actually reaches the provider; a configuration label alone is insufficient evidence. The Promptfoo configuration ties together the provider configurations, case inputs, and assertions. Pin the model identifier, supported reasoning setting, and relevant runtime versions with the results.

A simple layout makes those responsibilities visible:

evals/
  cases/          # Inputs, initial state, reviewed expectations
  providers/      # Adapters that invoke your application
  assertions/     # Deterministic checks and semantic rubrics
  configs/        # Role, handoff, and workflow comparisons
  results/        # Saved outputs, traces, grades, and usage
  decisions/      # Findings, review notes, and model choices

Start with one case and one configuration. Confirm that state resets between attempts, the intended agent receives the fixture, the trace contains its actual tool results, and an intentionally wrong answer fails the right assertion. Only then expand the case set and configuration matrix.

Save candidate outputs independently of their grades. Separate generation and grading stages are useful when calibrating a judge or comparing expensive agent episodes. You can then repair a deterministic check offline or recalibrate a judge against the same answers without regenerating every candidate. Preserve the original grades alongside any regrade.

For comparison views, use readable case descriptions as rows and clear model/reasoning labels as columns. Keep different roles and changed test contracts distinguishable so a chart does not imply an experiment you never ran.

Open a result cell's details to inspect Prompt & Output, then Evaluation for the assertion results. This makes the adapter's output inspectable alongside its grades.

Promptfoo Prompt and Output panel showing a case identifier and the saved answer with tool activity.

Promptfoo in practice. Our adapter receives a case ID and loads the fixture, so the Prompt field displays that ID instead of the full agent instructions. The saved output contains the answer and tool activity. Preserve the resolved inputs in the trace as well; the display label alone cannot reproduce an episode.

5. Use the simplest grader that can enforce the requirement

Match the grading method to the requirement. Exact answers and structured extraction may need equality or schema checks. Calculations can use executable checks with justified tolerances. Tool-using workflows need checks on actions, permissions, evidence, and limits.

Promptfoo's javascript assertion supports custom validation. Its llm-rubric assertion uses a model to assess written criteria. In our evaluation, we combined custom deterministic checks with a semantic judge and required both to pass. We fixed the judge at GPT-6 Sol with medium reasoning so changes in the candidate configuration did not also change the grader.

An episode is checked against executable rules and a semantic rubric, with both required for a complete pass.

Figure 3. One possible grading contract: both layers must pass. This illustrative episode passes semantic review but fails a required execution rule.

A semantic rubric should identify material errors, required coverage, acceptable uncertainty, and valid alternatives. Keep it scoped to the role and actual request. A specialist should not have to reproduce the coordinator's entire plan.

Decide how to combine the checks before comparing models. In our project, an answer passed semantic review but exceeded a declared tool-call limit. Its complete episode failed because the selection policy required compliance with that limit.

Actual Promptfoo assertion detail showing a failed workflow check and a passing semantic judgment.

Project example. The two graders reveal different properties of the same saved episode. A single overall score would hide the reason for its failure.

6. Calibrate the judge and investigate failures by layer

Before using scores to select models, test the grader on a small set of reviewed examples: clear passes, clear failures, and borderline cases. Keep the judge configuration fixed, inspect any disagreements, and determine whether the problem lies in the candidate answer, the reference label, or the rubric.

The judge also needs an accurate account of the candidate's evidence. Distinguish facts the agent received, facts it could retrieve, and facts available only to the evaluator. Missing a required retrieval can be an agent failure. Assuming the agent already saw a fact it never received can be a grading failure. Additional evaluator evidence can legitimately verify correctness or outcomes; the rubric must distinguish that use from claims about what the agent knew.

Two evidence paths from one source: a filtered tool result reaches the agent, while broader context reaches the judge.

Figure 4. A general example of unequal evidence. Check the grounding of a claim against what the agent received, and evaluate required retrieval separately. Broader evidence can still establish factual correctness or the final outcome; it should not be treated as something the agent already knew.

We encountered this in our five-agent application: a tool filter omitted relevant readings, while the judge received broader context. Investigating the trace exposed a tool defect and a grading assumption. Changing the candidate's reasoning setting would leave both defects in place.

Other failures exposed an incomplete definition of success. The application could recognize an intermediate recovery before the required outcome was complete. Similar mistakes are possible when a workflow treats “request submitted” as “task completed.”

For each failure, record the requirement, the supporting evidence, and the layer that needs attention: model, instructions, tool, orchestration, or grader. Codex is useful here because it can follow the trace into the implementation and reproduce many defects offline.

Be especially careful when a regrade improves scores. The same saved outputs went from 43 to 99 complete passes out of 135 after grading corrections. The answers were unchanged. This was evidence of a changed measurement, not improved model capability. Agreement on examples used to develop a judge also does not establish its accuracy on new cases. Our 27-example pilot was development calibration with Codex-reviewed labels. For higher-stakes decisions, have domain reviewers assess disagreements and check the frozen judge on separate cases.

In our regrade, a provider cache reused 27 earlier judgments despite a CLI-level cache setting. Those valid results saved cost, but they were not fresh observations of judge consistency. Record whether each answer and judgment was generated, retrieved locally, or billed with provider-side token caching; these are different forms of reuse.

What these problems mean for a new project

The investigation exposed several assumptions to address when building an evaluation from scratch:

What we encountered Real-world consideration Build this into the first eval
Recorded responses and stronger-model references still needed review A plausible answer can violate a refund policy, an engineering constraint, or a business rule Separate observed answers from reviewed expectations; document valid alternatives
A tool omitted facts that the judge could see Retrieval filters, permissions, or stale records can make an apparent reasoning failure an evidence problem Record the exact tool results received and tell the judge which facts were available
The simulator could score intermediate recovery as success Creating a ticket, submitting a job, or proposing an action does not establish completion Assert the required final state and any observation window, not just the agent's claim
Regrading unchanged answers changed the leaderboard The measurement can improve while the model stays exactly the same Calibrate against reviewed examples, version the rubric, and retain original grades
An answer passed semantic review but failed a tool-use limit Retries and extra lookups can violate latency, cost, or operational constraints Grade answer quality and workflow compliance separately; justify the limits before selection
Saved results were initially awkward to compare in the viewer Valid experiments can still produce a misleading report, especially across different test contracts Design readable case-by-configuration views early; label phases and preserve errors and untested cells

Treat a failed test as the start of a diagnosis. Use the trace to locate the defect before changing the model. Keep a findings log with the evidence, affected layer, correction, and whether the case or grading contract changed. This makes the evaluation useful for both improving the application and selecting models.

7. Harden the application when the eval exposes execution weaknesses

Building an eval can reveal that the application needs work before model selection is the main decision. Realistic episodes exercise retrieval, tool validation, handoffs, shared state, deadlines, and recovery. A leaderboard can look encouraging while those parts of the application remain unreliable. Treat the selected models as provisional engineering choices until the assembled system has demonstrated the behavior you require.

Make application hardening a standard part of the feedback loop in the evaluation process. A failed episode should lead to a diagnosis: did the model reason poorly, did a tool return incomplete evidence, did orchestration mishandle a failure, or did the grader measure the wrong thing? Promptfoo's tracing support helps inspect the actions behind a result. Budget time for this loop when the application has several agents, external tools, and stateful execution.

For example, our application concealed useful tool-validation details behind a generic error and propagated a specialist timeout without reliably preserving completed reports for review. We improved the feedback and made the failure policy explicit: a requested specialist that fails blocks the proposal, while completed reports remain available for diagnosis. Cancellation must be followed by checks to confirm that the remote work has stopped. Human authorization remains required for every executable proposal. Offline regression tests exercise these behaviors. They do not establish that a live agent will recover correctly or that the remaining evidence is sufficient for every decision.

This work can slow an initial model comparison because some apparent model failures turn out to be application defects. It also produces tests for weaknesses that happy-path demonstrations missed. In a support workflow, a specialist timeout might leave a refund policy unchecked; in an operations workflow, a completed subtask might be lost when another service fails. Define which missing assessments permit a limited proposal and which must block it. A stronger model does not automatically repair either contract.

Keep the loop bounded: reproduce the defect, fix the responsible layer, add an offline regression, and version the changed contract. Reuse saved outputs when investigating grading, but validate execution changes against the changed application before claiming improved reliability. Resolve defects that undermine measurements or required behavior; record lesser issues in the acceptance criteria. This is a stopping rule for a useful engineering phase, not a requirement to eliminate every imperfection before learning from an eval.

8. Compare configurations under a fixed decision policy

Define how quality, latency, and cost determine a choice. One product might prefer the cheapest configuration that meets the quality threshold. Another might prioritize complete success, then use latency to break ties. Record critical failures separately so an average cannot conceal them.

Start with a small pilot to verify the harness and observe duration and usage. An episode may contain several model and tool calls, followed by a judge call. Estimate the total workload from those observations rather than multiplying the number of test rows by the price per response. Report missing usage as unknown. Our long-running episodes made elapsed time a poor proxy for token spend. Set concurrency, timeouts, retry limits, and a spending boundary explicitly, then report duration and measured usage separately.

When an evaluation ends a session, stop the work, reconcile usage, save any unresolved accounting evidence, and then delete the session; deleting first can remove the evidence needed to finish accounting.

For a system with several agents, compare plausible configurations per role before exploring whole-team combinations. Exhaustively testing every combination quickly becomes expensive. Use the role results to narrow the choices, then validate the proposed team together. Agent behavior can vary even with identical inputs. Add repetitions appropriate to the decision's risk, report the variation, and keep failed attempts visible.

Actual Promptfoo view comparing three reasoning settings for one agent, with saved results and native charts.

Project example. Mission Director passed 7/9 cases with low reasoning and 8/9 with both medium and high reasoning. The charts summarize this role's saved results, not the reliability of the entire system.

Read down a column to find a configuration's weaknesses; read across a row to compare how configurations handled the same case. Use the failure filters and cell details to explain a difference before interpreting an aggregate score.

Promptfoo comparison of GPT-6 Sol and Astra at medium reasoning on the same nine Power cases.

Project example. A focused comparison keeps the case set and the level of reasoning fixed. GPT-6 Sol passed 8/9 complete episodes, and Astra passed 9/9. The visible failing cell identifies an evidence-lookup limit, connecting the model choice to a declared workflow requirement. These are saved final grades, not a fresh run.

We applied GPT-6 Luna at medium reasoning for Mission Director, Life Support, and Weather & Navigation, and GPT-6 Astra at medium for Power & Thermal and Risk Review. Our policy prioritized complete passes, then median episode time on ties, then lower reasoning effort. All five selections met an exploratory gate of at least eight complete passes out of nine in their respective role comparisons.

Nine related development cases per role supported provisional engineering choices, not statistical superiority or a validated, reliable system. The later execution findings require application hardening and separate validation; they do not retroactively change those saved comparison scores. Higher reasoning did not consistently improve results within these limits.

Verify access through the exact endpoint and credentials the harness uses. We initially confused a Codex-session model-access error with an error from the application API; those were separate paths. A tiny authorized smoke test can distinguish access problems from candidate failures before launching a matrix.

Set a stopping rule before the final comparison: resolve known measurement blockers, freeze the contract, run the agreed matrix, and make the decision. New observations belong in a findings log. They should lead to another experiment when they threaten the decision or reveal a meaningful product risk, rather than automatically expanding every run.

Check the assembled system and the path it takes

Evaluate the assembled system after selecting its components. Strong individual-agent scores do not establish that the workflow will succeed. Exercise real handoffs, shared deadlines, and recovery behavior on new cases before drawing conclusions about the team.

Grade outcomes and execution behavior separately. Check both whether the task was completed and whether the agents followed the required tool-use and authorization rules. A successful outcome can coexist with invalid tool calls or weak recovery; neither an outcome score nor a trace score tells the whole story.

Our mission follow-up illustrated both lessons: a specialist timeout stopped one assessment, while other missions reached their objectives despite rejected calculator calls. We kept those findings separate from the existing model rankings.

9. Give Codex a bounded evaluation task

Codex can help map agent boundaries, build recording and replay, draft cases, implement assertions, and trace failures back to their source. Its advantage is its ability to work across these connected artifacts. Keep product policy and consequential label disagreements subject to domain review; having the same assistant build and review a grader does not make that review independent.

A useful starting request is:

Inspect this multi-agent application and map its agents, handoffs, tools, and final outcomes. Propose success criteria and tests at the individual-agent, handoff, and complete-workflow levels. Build a small Promptfoo harness using the application's execution path, with isolated sessions and saved outputs. Keep observed responses separate from reviewed expectations. Add deterministic checks and a semantic rubric where needed, and document ambiguous labels or application defects. Prepare a pilot plan with estimated cost and duration for approval before paid execution. Preserve outputs so grading can be revised without rerunning candidates.

Review the workflow map, a few cases, and the grading rules before authorizing a broad run. This gives you concrete decisions to make early, when corrections are inexpensive.

The first milestone is a small evaluation package with explicit inputs, reviewed expectations, inspectable traces, and a documented selection rule. Add reviewed failures as regression cases, and rerun the relevant suites when prompts, tools, models, or policies change. Keep fresh validation cases separate from the examples used to make those repairs.

Source code

The repository contains the Mars Mission Control application and its Promptfoo evaluation harness, including test cases, the custom provider, deterministic checks, semantic grading rubrics, and model-selection documentation, so you can inspect the evaluation logic and adapt the workflow to your own multi-agent systems.

https://github.com/garystafford/openai-agents-api-mission-simulation-demo


This blog represents my viewpoints, not those of my past or current employers. All product names, images, logos, and brands are the property of their respective owners.

More from this blog

L

Latent Thoughts

2 posts

Practical, hands-on articles about AI engineering, cloud architecture, and software development. Through reproducible experiments, working code, and real-world examples, I explore how to build, evaluate, and operate AI-powered applications.