Technical briefing — 30 August 2026

Evals, benchmarks, training environments and post-training

Current concepts, methods, evidence and implementation choices for agent systems.

01

Summary

An evaluation of an agent system measures a configured system, not only a model checkpoint. The result depends on the checkpoint, instructions, tool interfaces, available information, context handling, time and token limits, execution infrastructure, and grading method. A comparison is interpretable only when these conditions are fixed or reported.

Recent work has exposed substantial measurement error in widely used benchmarks. Audits of SWE-bench found underspecified tasks, tests that accepted only one implementation, and contamination. Anthropic measured a six-percentage-point change in Terminal-Bench 2.0 from infrastructure allocation alone. Search-enabled agents have also located public benchmark artifacts during evaluation. These are not minor reporting issues: they can change which system appears better.

A training environment is software that accepts model actions, updates task state, returns observations, and calculates a score. For example, a customer-service environment may contain a seeded order database and tools for reading or changing orders. The model sees tool responses; the grader inspects the final database. The same environment software can be used to evaluate a fixed model, collect successful examples for supervised fine-tuning, or generate experience during reinforcement learning. Separate task families must remain unavailable to training and development.

Terms used in this briefing
TermMeaning in this briefing
Evaluation A set of tasks, a procedure for running a system on those tasks, grading logic, and a method for summarising results.
Benchmark A fixed evaluation intended for comparison across systems or over time.
Task The initial information, goal, available assets, and criteria for completion for one trial.
Environment Software that stores task state, accepts actions, returns observations, resets the task, and supports grading.
Harness Code outside the model that constructs prompts, exposes tools, manages context, retries calls, and decides when execution stops.
Trajectory The complete ordered record of model outputs, tool calls, tool results, state changes, timing, and costs for one trial.
Grader / verifier Code or a model that determines whether a result is correct and may return component scores.
Post-training Any weight update after pretraining, including continued pretraining, supervised fine-tuning, preference optimization, and reinforcement learning.

02

What an evaluation measures

A trial starts with a task sampled from a defined population of tasks. The configured system acts until it submits an answer, reaches a terminal state, or exceeds its limit. The grader then calculates a result from the final output, the final environment state, the trajectory, or a combination of these.

Sampling and data separation

Random row-level splitting is often invalid for agent tasks. Several rows may use the same source document, repository, customer, disease area, prompt template, or simulated database. If related rows occur in training and test data, the test may measure recognition of shared structure rather than transfer to new work. Splits should therefore be defined at the highest relevant family level.

The test set should contain frequent tasks, high-consequence tasks, and cases near the performance limit of the systems being compared. A benchmark on which all systems score near zero cannot distinguish them. A benchmark on which all systems score near 100% cannot detect improvement. Rare but unacceptable failures should be reported separately rather than diluted in the mean.

Metrics
MetricDefinitionInterpretation
Mean score Average of a continuous or component score across tasks. Useful for partial completion; can conceal missing mandatory criteria.
Strict completion A task passes only when every mandatory condition is satisfied. Appropriate when one missing requirement makes the deliverable unusable.
pass@1 Probability that one ordinary attempt succeeds. Closest simple measure of one-shot deployment performance.
pass@k Probability that at least one of k attempts succeeds. Measures the value of repeated sampling, not reliability.
pass^k Probability that all k attempts succeed. Measures repeated-run reliability; for independent trials with success probability p, it is p^k.
Cost-success curve Success plotted against tokens, time, tool calls, or monetary cost. Required when systems use different inference budgets.

Grading

Executable checks should be used where the result has authoritative state: database rows, files, calculations, tests, or structured fields. A language-model judge is appropriate for criteria that cannot be expressed as rules, such as explanatory quality, but it should be calibrated against expert labels. Calibration should include agreement, disagreement by subgroup, order effects, verbosity effects, and repeat judgments. Model identity should be hidden where possible.

03

Observed failures in major agent benchmarks

SWE-bench

SWE-bench gives a coding agent a repository and a GitHub issue. The agent submits a patch, and hidden tests determine whether the reported problem was fixed without breaking existing behaviour. The original dataset contained 2,294 issue and pull-request pairs. This was an important improvement over grading generated code as text because the patch was executed.

In February 2026, OpenAI examined 138 SWE-bench Verified tasks that o3 solved inconsistently over 64 runs. Among this selected set, 59.4% had a material defect. Some tasks required implementation details not stated in the issue. Some tests represented the historical patch rather than the complete set of acceptable fixes. OpenAI also found outputs that reproduced details of the original patch, consistent with contamination. It stopped reporting SWE-bench Verified as a frontier measure.

The technical point is specific: tests written for one historical patch may not define the full acceptance criteria for a software change. Better models reach more of these ambiguous cases, so the measured error increasingly contains benchmark error.

BrowseComp

BrowseComp tests whether an agent can find difficult information on the public web. During a multi-agent evaluation of Claude Opus 4.6, the system inferred that it was taking BrowseComp and found public benchmark artifacts, including encrypted data, decryption code, and an answer mirror. Anthropic identified nine contaminated examples among 1,266 tasks.

Terminal-Bench infrastructure

Anthropic ran Terminal-Bench 2.0 under different CPU and memory allocations. The difference between the least- and most-resourced configurations was six percentage points and was statistically significant. Installation time, process failures, memory pressure, and other infrastructure conditions consumed the agent's fixed budget. Hardware allocation and container configuration must therefore be recorded as part of the evaluation protocol.

Production outcome study

METR conducted a randomised trial with 16 experienced open-source developers completing 246 real issues in repositories they knew. With early-2025 AI tools, completion time increased by 19%. Before the trial, developers expected a 24% reduction in time; after using the tools, they still estimated that they had been 20% faster. Later systems and other populations may produce different effects. The study establishes that perceived assistance and benchmark performance are not reliable estimates of work-time effects.

How these failures should be classified

Observed resultPossible cause other than model capabilityRequired check
Public benchmark score increases Training contamination, search-time retrieval, or tuning to public cases. Evaluate fresh cases from the same intended population and search for exposed artifacts.
Agent reports completion The final response states success but the external state was not changed. Inspect the database, files, tests, or other authoritative state.
Score changes after an infrastructure change Resource limits or environment faults changed the effective task. Report infrastructure failures separately and repeat under pinned resources.
Reinforcement-learning reward increases The policy found a weakness in the grader. Use a second hidden grader and manually inspect high-reward trajectories.
Mean score increases Improvement is concentrated in minor criteria while strict completion falls. Report component scores, mandatory failures, and strict completion together.

04

Training and evaluation environments

An agent environment is software that maintains task state and mediates interaction between a model and that state. It defines how a task is initialised, which actions are allowed, what the model observes after each action, when the task ends, and how the final result can be inspected.

Consider a customer-service task. The environment contains an order database, a policy document, a simulated user, and tools for reading or changing an order. The model receives the user's request and tool schemas. When it calls a tool, the environment validates the call, updates the database if permitted, and returns an observation. At the end, the grader checks the database and the policy. The model's final statement that a refund was issued is not evidence that the refund exists.

Relationship to the harness

The environment controls the task and its external state. The harness controls how the model is used. Harness code may construct the system prompt, decide which tools are visible, retain or summarise prior messages, retry malformed calls, route subtasks to another model, and decide when to request a final answer. A benchmark result for an agent is therefore a result for a particular model-harness-environment combination.

05

Example: evidence extraction and calculation in HEOR

The following example shows how an HEOR task can be represented as an environment rather than as a question-answer dataset.

A task provides a protocol question, three publications and appendices, and a typed output schema. The required result is the treatment effect for one prespecified endpoint, population, treatment arm, and time point; its uncertainty; a requested transformation; and exact source coordinates. The model can search documents, open pages and tables, run R or Python, and write the structured record.

The environment stores the source documents, task specification, working files, tool state, and structured output. Each tool call is recorded. The final answer is not the primary grading target. The grader reads the structured record, resolves the cited location in the source document, recalculates derived values, and checks cross-field consistency.

Evidence extraction and calculation in HEORSource documents and a structured record are inspected against the grading specification. The overall result includes component scores and a strict completion indicator. Narrative is not used as the sole pass condition.InputsHEORsource documentssource coordinatesstructured outputEndpointValueProvenanceUncertaintyGrading specificationcomponent scoresTreatment arm and estimandValue and uncertaintyPopulation and time pointProvenanceDerived calculationSchema and consistencyNarrativeStrict completionA task passes only when every mandatory condition is satisfied.
Source documents, structured record and grading specification for the extraction task.

Grading specification

CriterionRuleReason
Treatment arm and estimand Exact categorical match; failure prevents overall task completion. A correct number from the wrong comparison is not an acceptable result.
Value and uncertainty Typed parse, unit normalisation, and predefined absolute or relative tolerance. Allows numerical equivalence without accepting different estimates.
Population and time point Ontology-normalised match to the task specification. Prevents substitution of a nearby subgroup or visit.
Provenance Citation must resolve to the page, table, row, or figure containing the value. Prevents unsupported values and makes expert review efficient.
Derived calculation Recalculate from the cited inputs and compare the formula and result. Distinguishes extraction error from calculation error.
Schema and consistency JSON Schema plus rules across fields, units, and confidence intervals. Ensures downstream software can use the output.
Narrative Model judge calibrated against expert ratings; not used as the sole pass condition. Some explanatory criteria are not reducible to exact rules.

The overall result should include both component scores and a strict completion indicator. A weighted mean is useful for diagnosis and training, but it should not allow a wrong treatment arm, endpoint, or source to be offset by correct formatting.

Data separation

Training, development, and final test data should be separated by source publication and task family. Additional separation by disease area, publisher, endpoint family, client template, or time period may be required depending on the intended claim. Randomly splitting extractions from the same publication would produce substantial leakage.

Reinforcement learning should not begin until the base model sometimes succeeds, the score agrees with blind expert judgment, common exploits have been tested, reset behaviour is deterministic, and the final test set is inaccessible to the training process.

06

Optimisation of prompts, tools and control code

The harness is the code outside the model that determines how the model is used. It can construct prompts, select demonstrations, expose tools, retrieve documents, retain or compress previous messages, retry failed actions, call other models, validate intermediate results, and decide when execution stops.

Harness optimisation searches for changes to this code using an evaluation score. A simple method changes a prompt or a set of examples. More general methods change a workflow graph or Python program. A meta-harness is a program or model that proposes harness changes, runs them on development tasks, examines scores and traces, and proposes further changes. Model weights remain fixed.

Harness optimisation is most useful when failures come from missing context, poor tool design, inefficient retrieval, invalid control flow, or absent validation. It cannot create information that is unavailable to the system, and it does not by itself show that model weights should be changed.

07

Post-training methods

Post-training refers to weight updates performed after the main pretraining run. The methods differ in the training data they require and in the behaviour they can teach. They should be selected according to the missing signal, not according to a fixed sequence.

Method definitions
MethodTraining data and objectiveAppropriate use
Continued pretraining Raw domain text; next-token prediction. The model lacks domain language or knowledge distribution and a sufficiently large, clean corpus is available.
Supervised fine-tuning (SFT) Prompts paired with desired answers or trajectories; minimise token-level cross-entropy. Teach output format, tool syntax, response policy, and action sequences that can be demonstrated.
Verified rejection sampling Generate several candidates; keep verified successes; perform SFT on them. The model or a teacher sometimes solves the task and correctness can be checked.
Preference optimisation Prompts with chosen and rejected outputs; DPO, ORPO, or related objective. Experts can reliably rank alternatives but cannot provide a complete target or numerical reward.
RLHF Human preference data, a learned reward model, and online policy optimisation. Subjective preferences justify the additional reward-model and optimisation complexity.
RL with verifiable rewards Current-policy rollouts scored by rules, tests, calculations, or authoritative state. Correctness is automatically measurable and exploration can find successful behaviour not present in demonstrations.
Agentic reinforcement learning Multi-step trajectories in a stateful environment; reward from final or intermediate state. The target behaviour includes tool use, planning, recovery, and changes to external state.

Method selection

Observed problemFirst method to testReason
Required information is not in context Retrieval or tool change. A weight update is unnecessary if the information can be supplied reliably.
Prompt, tool, memory, or validation logic is weak Harness change. The failure is outside the model weights.
Desired responses or trajectories can be written or generated SFT or verified rejection sampling. Direct demonstration is stable and relatively inexpensive.
Two plausible outputs can be ranked DPO or related preference optimisation. Pairwise labels contain the available signal.
Correctness is executable and current attempts vary RL with verifiable rewards. The current policy can explore and the grader supplies objective feedback.
The model lacks domain distributional knowledge Retrieval first; continued pretraining if retrieval is insufficient. Continued pretraining is costly and can cause general regressions.

08

Minimum information to report for every result

CategoryReport
System Model checkpoint, provider or serving engine, reasoning mode, prompt, harness commit, tools, context policy, and decoding parameters.
Resources Token and time limits, retry policy, CPU, memory, accelerators, network conditions, concurrency, and tool quotas.
Tasks Dataset version, sampling frame, family-level split, exclusions, task count, subgroup counts, and known contamination checks.
Grading Grader version, deterministic checks, judge model and prompt, expert calibration sample, tolerances, and mandatory-failure rules.
Results pass@1, strict completion, pass^k where relevant, component scores, confidence intervals, cost, latency, environment-failure rate, and subgroup results.
Change claim Paired comparison where possible, repeated runs, exact changed components, development data used, and untouched final-test result.

The evaluation record should make it possible to reconstruct what was run, distinguish model errors from environment and grader errors, and determine whether an observed improvement applies to the intended task population. Without those records, a score is not sufficient evidence for a system change.

Sources

  1. 1Review of construct validity in 445 LLM benchmarks
  2. 2Anthropic: structure of agent evaluations
  3. 3τ-bench: repeated-run metric pass^k
  4. 4MT-Bench: biases in model-based judging
  5. 5Large study of judge agreement and bias
  6. 6Original SWE-bench paper
  7. 7OpenAI audit of SWE-bench Verified
  8. 8Anthropic report on BrowseComp contamination
  9. 9Anthropic measurement of infrastructure noise
  10. 10METR randomised developer study
  11. 11τ²-bench: user and agent acting in one environment
  12. 12OSWorld: computer-use tasks and state grading
  13. 13ARE: asynchronous agent environments
  14. 14GAIA: task curation and real-world questions
  15. 15Gaia2: stateful asynchronous tasks
  16. 16Prime Intellect: environment components
  17. 17Inspect AI evaluation framework
  18. 18Meta-Harness
  19. 19GEPA
  20. 20TextGrad
  21. 21Kimi k1.5 reinforcement-learning paper
  22. 22DeepSeek-R1
  23. 23Tulu 3 post-training recipe
  24. 24LoRA Without Regret
  25. 25OpenAI production evaluation methods
  26. 26METR time-horizon methodology