Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

inspect_evals tasks

Classic eval benchmarks from the inspect_evals catalog — plus the bundled keyless hello smoke test — one experiment per task.

Each task is its own experiment; there is no task parameter. The experiment list is generated from the pinned upstream catalog (about 180 experiments). A few examples:

ExperimentTaskNeeds
inspect-hellobundled hello (2 instruction-following samples)any model — your real one, or keyless/offline with mockllm/model; both deterministically score 1.0
inspect-gsm8kinspect_evals/gsm8k (grade-school math)a provider + network
inspect-gpqa-diamondinspect_evals/gpqa_diamond (graduate-level science MCQ)a provider + network
inspect-swe-bench-verified-miniagentic tasks whose eval declares a sandboxthe above + docker on the host

Tasks whose eval needs an extra pip dependency (gaia, agentdojo, …) are catalogued but not yet runnable. The full list is the overview page in the web GUI, or nix flake show.

$ nix run .#inspect-hello -- \
    --set model=mockllm/model \
    --set limit=0 \
    --set epochs=1 \
    --set 'generate_args={}'

$ nix run .#inspect-gsm8k -- \
    --set model=anthropic/claude-sonnet-4-5-20250929 \
    --set limit=20 \
    --set epochs=1 \
    --set fewshot=10 \
    --set fewshot_seed=42 \
    --set shuffle_fewshot=true \
    --set 'generate_args={}'

Every param is on the command line — experiments have no defaults. Copy commands from the composer, or run bare and copy the suggested command it prints.

Real tasks download their datasets from HuggingFace on first run. inspect-hello bundles its samples and stays offline — it exists so your first run (and any CI check) exercises the whole pipeline with zero setup.

Parameters

A task’s own arguments are real typed params, taken from the task function’s signature — inspect-gsm8k has fewshot, fewshot_seed, shuffle_fewshot; other tasks have their own. Upstream defaults prefill the form, and a kwarg whose upstream default is None becomes a nullable param bound explicitly as --set name=null — the oneliner always states every value, so an upstream default change can never silently reinterpret it. A task kwarg that shares a family param’s name (seed, limit, …) appears prefixed, as task_seed etc.

The form lists the task’s params first (they’re the condition’s substance), then model, then the inspect harness knobs:

ParamTypeForm prefillNotes
per-task paramstypedthe task’s own defaultse.g. fewshot on inspect-gsm8k; null where upstream declares None.
modelllmmockllm/model for inspect-hello; none elsewhereInspect model id (openai/…, anthropic/…); the prefix also keys credential injection. Real tasks have no prefilled model — the canonical condition is always something you chose.
limitint0Sample cap; 0 = the whole dataset.
epochsint1Passes over the dataset.
generate_argsobject{}Generation overrides, rendered as a typed form (temperature, max_tokens, reasoning_effort, …). A field left unset keeps the provider’s default.

Results

The standard inspect-family results: status, samples, completed, errors, score, score_name, tokens_input, tokens_output.

Identity

The experiment name is part of the condition hash, so tasks never collide. All the tasks share one pinned upstream package, so updating that pin re-versions the family together. Every run records the exact inspect_ai/inspect_evals versions, task version, and dataset identity, so buckets that a version boundary didn’t really change can be pooled at read time — and genuine breaks flagged — later. See ImpossibleBench for how families sit on the shared wrapper.