Skip to main content

Command Palette

Search for a command to run...

Choosing the Right OpenAI Model and Reasoning Level

An evaluation guide to accuracy, latency, and cost across complex workloads.

Updated
•17 min read•View as Markdown
Choosing the Right OpenAI Model and Reasoning Level
G
Applied AI Engineer at OpenAI. I write about building AI applications, evaluating models, and turning experiments into practical solutions, drawing on my background in cloud architecture and software engineering.

Your AI application needs answers it can trust, delivered fast enough and at a cost that makes sense. Choosing the right model and reasoning level means defining those requirements and testing which configurations meet them.

In my earlier post, I compared models across multiple providers. In this post, I narrow the eval to OpenAI’s GPT-6 Astra, Sol, and Luna at their supported reasoning settings, using Promptfoo to examine how trade-offs vary across workloads. As the latest reasoning models handled earlier tasks with ease, I increased task complexity and broadened coverage across more domains to better reveal differences in their capabilities.

Define what success means

Evaluate against the requirements of your application:

Reason to evaluate Question to answer
Accuracy Does it get your tasks right?
Latency Is it fast enough for the application?
Cost Is the quality worth the cost per task?
Consistency Does it succeed across repeated attempts?
Workload fit Which model and reasoning level works best for each use case?

This post focuses on accuracy, latency, and cost, using repeated attempts and different domains to examine consistency and workload fit.

An interactive application may need a quick answer; a background planning job may take longer to produce a better plan. A small error may be acceptable in an estimate and unacceptable in a release gate.

A key part of defining requirements is deciding what counts as correct and making sure the grader measures it fairly.

Does a correct schedule need to be feasible or optimal? What units and rounding rules apply? Does list order matter? Define these requirements before looking at model scores.

Define accuracy, latency, and cost requirements

Figure 1. Keep the requirements short enough to use and precise enough to grade.

Specify whether the threshold applies overall or to every task. Our default gate requires at least 95% observed accuracy overall and per task, with no critical violations. With three attempts per task, this requires 3/3: 2/3 falls below 95%. A configuration scoring 80/81 overall (98.8%) can therefore fail the per-task gate. An overall-only gate admits different configurations without changing their scores.

Use representative tasks

The benchmark contains 27 distinct complex tasks across nine domains, three tasks per domain. Each provides a fictional source pack with the necessary facts and rules; the GPT-6 sweep (one complete evaluation run across the benchmark matrix) uses no external search or tools.

These domains test different reasoning skills: optimization under constraints, reconciling calculations with authoritative records, resource allocation, causal explanation, and safe sequencing. They let us compare how the same configuration handles different business and engineering workloads.

Strong early performance shaped the experiment. Every Astra setting solved every attempt in the earlier mixed-difficulty benchmark. I expanded the task set and increased its reasoning complexity by adding interacting constraints, multiple steps, and explicit optimality or safety requirements. Simpler cases remain as regression tests. As models improve, familiar questions may no longer reveal differences.

Domain What the tasks exercise
Operations planning Fulfillment optimization, production scheduling, warehouse queues
Procurement Sourcing constraints, contingent purchases, invoice matching
Liquidity Cash ledgers, currency netting, refinancing covenants
Financial diligence EBITDA reconciliation, revenue cohorts, cash earnings
Transformation Roadmap dependencies, impact estimates, automation policy
Engagement delivery Skills and capacity, work assignments, rescue plans
Incident response Safe retries, fault explanations, lease fencing
Disaster recovery Backup chains, consistent restore points, resilient capacity
Release engineering Canary analysis, migration sequences, version compatibility

These are different task types, not numerical variants of a single template. Difficulty is a design judgment. Expected answers are frozen, checked against reference solvers and source data, and tested with valid and invalid grader inputs; experts have not reviewed them independently.

Example: How much of this invoice should we pay?

One task asks the model to reconcile an invoice with a purchase order, delivery records, returns, and credits. The records are summarized below; each line is evaluated separately.

Line Order and agreed price Received / returned Invoice Credit
P1 10 X at $100 each 6 X + 8 Y received; 2 Y returned 10 X at $100 $50
P2 8 X at $120 each 8 X received; 1 X returned 8 X at $125 None
P3 5 Z at $80 each 4 Z received 5 Z at $80 None

A signed agreement allows two Y units to replace one X. An unsigned draft proposes a one-for-one substitution. Receipt R2 (the eight Y units) and credit C1 ($50) each appear twice in the source pack.

The abbreviated prompt captures the decision rules:

Determine how much to pay on each invoice line.
Count duplicate receipt and credit IDs once; subtract returns.
Use the signed substitution ratio. Ignore the unsigned draft.
Do not move surplus quantities between purchase-order lines.

Quantity matched = minimum of net received, invoiced, and ordered
quantities, expressed in equivalent ordered units.
Held quantity = invoiced quantity minus quantity matched.
Held quantity measures shortages, even when the price is disputed.

If the invoice price differs from the agreed price, pay $0
for that entire line and report price_dispute.
Otherwise pay quantity matched × agreed price, subtract posted
credits once, and floor at $0. Report short_receipt or matched.

Return JSON with payable_dollars, held_equivalents, reasons,
total_payable_dollars, deduplicated_record_ids, and evidence_ids.
Sort ID lists lexicographically; cite authoritative evidence only.

The checked answer is $1,170, payable as follows: $850 for P1, $0 for the price-disputed P2, and $320 for P3. Getting there requires more than arithmetic: the model must deduplicate records, convert substitutions, distinguish shortages from price disputes, and follow the signed evidence. See the complete prompt and source pack (invoice_match).

Make the comparison reproducible with Promptfoo

The repository was converted to Promptfoo to bring execution, grading, and inspection into one workflow. Promptfoo runs the configuration matrix, applies assertions, and saves the responses and metrics.

Each viewer row is a task attempt; each column is a model/reasoning configuration. Headers summarize scores, tokens, cost, and average latency. Use Columns to focus a comparison, Failures to filter strict failures, and cell details to inspect prompts, outputs, and grading.

Promptfoo results interface with two model configurations selected

Figure 2. The combined view has 17 providers and 81 task-attempt rows. The selected columns compare Sol/low and Luna/max; the green percentage shows the strict pass rate. Column selection changes the view, not the dataset. View full size.

A repository adapter delegates transport to Promptfoo’s native OpenAI Responses provider, supplies the task-specific schema, calculates usage-based cost, and selects the final assistant answer for grading:

Evaluation workflow using Promptfoo and the Responses API

Figure 3. Shared source packs and declared contracts flow through the same evaluation harness for every configuration.

Every request uses OpenAI Structured Outputs: text.format with JSON Schema and strict: true. Schemas define field types and input identifiers, without embedding solved answers.

A simplified request fragment illustrates the contract:

const requestFragment = {
  model: "gpt-6-sol",
  reasoning: { effort: "low" },
  text: {
    format: {
      type: "json_schema",
      name: "payment_summary",
      strict: true,
      schema: {
        type: "object",
        properties: { total_payable_dollars: { type: "integer" } },
        required: ["total_payable_dollars"],
        additionalProperties: false
      }
    }
  }
};

Each task has its own schema. Deterministic Python assertions check arithmetic, feasibility, optimality, evidence, and unsupported conclusions.

Task accuracy measures the decision; strict correctness also requires compliance with the output contract. Unordered selections compare as sets for accuracy, but must satisfy sorting rules for a strict pass. Execution sequences stay ordered. A critical violation breaks a hard constraint, such as a scheduling or recovery rule; a feasible but suboptimal plan can fail accuracy without one.

The matrix uses 17 configurations: five Astra efforts (low through max) and six each for Sol and Luna (none through max). The model IDs are gpt-6-astra, gpt-6-sol, and gpt-6-luna. Check the current model guidance before adapting the comparison to another release.

Eval actions → View YAML shows the run’s saved provider settings:

Saved Promptfoo configuration showing the Astra low provider settings

Figure 4. The Astra/low provider entry shows the API endpoint, service tier, and output-token budget. View full size.

Preview the scope, run one attempt per task/configuration, then inspect and report the saved results:

npm ci
npm test
# Preview without API calls
BENCHMARK_RUNS=1 BENCHMARK_CONCURRENCY=4 npm run eval:matrix:plan
# Run 459 task/configuration pairs
BENCHMARK_RUNS=1 BENCHMARK_CONCURRENCY=4 \
  npm run eval:matrix -- --env-file .env.local
npm run eval:view -- --port 15500
npm run eval:report -- results/promptfoo-YOUR-RUN.json

The ignored env file supplies OPENAI_API_KEY. I ran three unchanged sweeps: 27 tasks × 17 configurations × three attempts = 1,377 evaluated attempts. Requests use the Standard service tier, with a 16,384-output-token budget, default verbosity, and a concurrency of 4. Promptfoo response caching is disabled; API input-prefix caching can still occur. Source hashes verify matching benchmark definitions across sweeps.

To run three attempts in one evaluation, set BENCHMARK_RUNS=3. The article reproduction guide also documents combining separate sweeps as done here.

Inspect the completed evaluations

Promptfoo’s headline passing percentage reflects strict grading. The article’s accuracy charts use task_success, so a sorting-only failure does not become a wrong business decision.

Click Show Charts to view summaries of strict pass rates and named scores before inspecting individual answers.

Promptfoo Show Charts view with native pass-rate and named-score charts

Figure 5. Promptfoo’s native charts summarize the three-sweep results. The publication charts below use median latency; the viewer displays average latency. View full size.

The report checks coverage, distinct uncached responses, comparable settings, and critical violations. Use --accuracy-gate overall --min-accuracy 0.95 when your requirement is aggregate accuracy.

Latency is provider-call duration, not total matrix runtime. We report medians; applications with deadlines may also need p95 latency.

Why repeat the same tasks? A configuration can answer correctly once and miss a requirement on its next attempt. Repeats expose variation and provide cost and latency summaries with a larger sample, without adding new benchmark questions.

Compare accuracy, latency, and cost

Across the sweeps, 1,234 of 1,377 attempts passed task accuracy; 1,229 passed strict grading. Completed answers had no schema violations or final-answer extraction failures.

The known-usage estimated cost subtotal was $19.54, including failed answers. One Luna/none API refusal supplied no answer or usage. It counts as an unsuccessful, ungraded attempt with unknown cost; Luna/none is omitted from cost comparisons.

There were 29 critical violations: 28 in none configurations and one in Luna/low. Valid structure did not prevent substantive errors or the five strict-only failures.

Accuracy for all 17 model and reasoning configurations

Figure 6. Each configuration has 81 attempts on 27 distinct tasks. Bars count successful task attempts; markers count strict passes. The hatched segment identifies the ungraded API refusal. View full size.

Astra/low is a strong default in this benchmark. It achieved 81/81 task successes and strict passes, with no critical violations. Higher Astra efforts also reached 81/81, at greater cost and latency. If Astra/low meets your response-time and budget requirements, these results give no accuracy-based reason to increase its effort.

Under the default per-task gate, Luna/max was the least expensive qualifying configuration; Sol/high had the lowest median latency.

Configuration Correct responses Strict passes Mean estimated cost/attempt Median latency
GPT-6 Luna / max 81/81 81/81 $0.001505 13.889 s
GPT-6 Sol / high 81/81 81/81 $0.008646 5.015 s
GPT-6 Astra / low 81/81 81/81 $0.024565 5.265 s
GPT-6 Sol / low 78/81 77/81 $0.005723 3.722 s

All four had zero observed critical violations. The first three meet the default per-task gate. Sol/low meets the overall-only 95% gate but misses the per-task requirement.

Compared with Astra/low, Sol/high reduced mean estimated cost per attempt by about 65%, with all 81 strict passes. Its median latency was slightly lower: 5.015 versus 5.265 seconds. The small latency difference does not establish a consistent speed advantage.

Luna/max reduced mean estimated cost per attempt by about 94% relative to Astra/low, also with all 81 strict passes, but took about 2.6 times as long at the median. That trade-off may suit a background job with a flexible deadline.

Repeats changed the shortlist. Sol/low scored 27/27, 26/27, and 25/27. Its later regional-failover plans had modeled costs of $290 and $350 against a $260 optimum, and it returned a suboptimal fulfillment plan. The first sweep alone missed those errors.

Sol/low reduced mean estimated cost by about 77% and median latency by about 29% relative to Astra/low. Whether those savings justify its observed misses depends on the accuracy requirement.

Accuracy plotted against estimated cost by model

Figure 7. Outlined points mark each model’s least expensive qualifying effort. Panels use separate linear cost scales; compare dollar labels. Luna/none’s cost is unavailable. View full size.

Accuracy plotted against median latency by model

Figure 8. Outlined points mark each model’s fastest qualifying effort. Latency panels use separate scales. View full size.

In the combined view, higher points indicate greater accuracy, points farther left indicate greater speed, and smaller bubbles indicate lower cost.

Accuracy and latency with bubble area representing mean estimated cost

Figure 9. The bubble area represents the mean estimated cost per attempt; all models share the same accuracy and latency scales. Luna/none is omitted because its cost is incomplete. The vertical scale begins at 40%. View full size.

With an overall threshold of only 95%, Luna/high becomes the least expensive qualifying choice: 77/81 (95.1%), $0.000774 per attempt, and a median of 7.979 seconds. Sol/low becomes the fastest. Both fail the per-task gate. These are different selection rules applied to the same results.

Costs are based on recorded usage and the repository’s frozen September 28 rate card. API input-prefix caching differed across sweeps and affected costs; disabling response caching does not disable prefix caching. These estimates are not billed totals. Preserve usage and rates with the results, and check current pricing for your workload.

A small GPT-5.6 comparison

GPT-6 Sol/high matched GPT-5.6 Sol/high’s accuracy, while costing 59% less and delivering 30% lower median latency. Both passed all 27 tasks. The mean estimated cost per task fell from $0.02331 to $0.00953, and the median latency fell from 7.475 to 5.264 seconds. GPT-5.6 Sol/low scored 25/27, missing a fulfillment optimum and an incident attempt count.

GPT-5.6 and GPT-6 Sol compared on accuracy, latency, and cost

Figure 10. One attempt on the same 27 tasks, using identical prompts, schemas, graders, and run settings. GPT-6 values use the first sweep, keeping this comparison separate from the three-attempt results. View full size.

These observations support testing a newer model against your existing workload. One attempt per task does not ensure reliability; separate runtimes and input-prefix caching can affect latency and cost.

Match the configuration to the domain

The overall score hides useful differences. Luna/low passed 67 of 81 attempts, yet stayed at 9/9 in transformation, engagement delivery, disaster recovery, and release engineering. For those domain samples, it was the least expensive qualifying choice.

Task accuracy by domain and configuration

Figure 11. Every cell shows the correct responses across nine attempts: three distinct tasks, three attempts each. A 9/9 cell is an observation in this sample, not a domain reliability estimate. View full size.

Sol/low was the fastest among qualifying configurations across seven domain samples; Astra/low led in operations planning, and Sol/medium in disaster recovery. Repeats changed the procurement cost choice from Luna/medium in the first sweep to Luna/max across all three.

Least expensive and fastest qualifying choices by domain

Figure 12. Mean estimated cost and median latency for the selected domain configurations. Each selection requires all nine responses to succeed, with no critical violations observed. View full size.

The qualifying choices differ by domain under the same selection rule. Use them as a shortlist for your own tasks.

Inspect failures that change the decision

Structured Outputs enforces structure; the grader still has to measure the right decision. Two responses from the first sweep illustrate the distinction.

Correct plan, wrong ordering. Sol/low returned the optimal $1,090 fulfillment plan. These fields are excerpted from its answer:

{
  "selected_orders": ["O1", "O3", "O5", "O6", "O8", "O10"],
  "net_value_dollars": 1090
}

Alphabetical ordering means lexicographic string order: O10 precedes O3. The schema validates array structure but does not impose ordering. The grader’s named scores distinguish the correct decision from the strict failure:

{
  "task_success": 1,
  "format_compliance": 0,
  "critical_violation": 0,
  "strict_correctness": 0
}

Feasible plan, lower value. Luna/medium returned a feasible $1,080 fulfillment plan against the $1,090 optimum. It passed format compliance and had no critical violations, but failed task accuracy because the task requires the optimum. An application could instead define an acceptable dollar shortfall before running the evaluation.

Ambiguous business meaning. In the invoice task, a purchase-order line has eight invoiced equivalents and seven quantity-matched equivalents. Its invoice price also differs from the signed price. The required output includes:

{
  "held_equivalents": { "P2": 1 },
  "payable_dollars": { "P2": 0 },
  "reasons": { "P2": "price_dispute" }
}

The held equivalent is the quantity shortfall; the price dispute separately blocks payment. We clarified that distinction before these sweeps, without changing the expected answer. All 51 responses returned the intended held quantity and duplicate-ID list; 48 passed the full task.

Selecting the completed answer. OpenAI’s message-phase guidance distinguishes commentary from final_answer. Luna/none returned both on the fault-hypothesis task. The adapter preserved the raw messages and graded only the final answer, which passed.

Task correctness and strict-only failures across the sweep

Figure 13. The 1,377 attempts’ separate strict passes, strict-only failures, substantive mistakes, and an ungraded API refusal. Critical violations are tracked separately.

Apply the method to your workload

Start with representative application tasks and agreed scoring rules. Check the grader with answers that should pass and those that should fail. Compare configurations using the same prompts and output contracts, repeat the attempts, and inspect the misses before choosing.

Select a configuration using accuracy, latency, and cost requirements

Figure 14. Apply requirements first, then compare the eligible configurations.

Three attempts cannot establish production reliability, and three tasks cannot represent an entire domain. This benchmark omits retrieval, tools, long conversations, and operational conditions. Validate your shortlist on additional application inputs; preserve the task definitions, scoring rules, model IDs, and run conditions.

A brief GPT-6.1 agent experiment

I also tried GPT-6.1 Sol at medium reasoning on diligence reconciliation, with hosted multi-agent support enabled. The first call produced no subagents. A second call explicitly requested three independent checks: closing entries and EBITDA; addbacks, debt, leverage, and valuation; and source authority and missing evidence.

The second call spawned closing_check, financial_check, and authority_check. The response exposes their names and the actions they spawned. Inter-agent messages are encrypted, so the diagram shows requested scopes, not their conversations.

Observed GPT-6.1 delegation with requested independent-check scopes

Figure 15. Three subagents observed after explicit delegation. Enabling the capability alone did not produce any subagents in the other call.

Both final answers passed the same task grader. The explicitly delegated call took longer and had a higher estimated cost:

Two exploratory GPT-6.1 diligence calls

Figure 16. Two observations on a single task, under different prompts and caching conditions. This is not a controlled estimate of delegation overhead. View full size.

A future post will explore when delegation improves quality enough to justify its additional latency and cost.

Source code

The repository contains the Promptfoo configuration, source packs, output contracts, answer keys, reference solvers, deterministic graders, report generator, and figure script.

https://github.com/garystafford/openai-reasoning-benchmark


This blog represents my viewpoints, not those of my past or current employers. All product names, images, logos, and brands are the property of their respective owners.