What an AI coding-agent run will cost: what can be known before it starts, and what can only be seen while it runs

Code, pre-registration and outputs: https://github.com/vhsgreed/cost-predictor-spike (article version cited at commit e9ce211).

Version: 1.0, 2026-10-07. Changes from the first draft are listed at the end. Author: Karl Sundström (Agent Recourse), corresponding and accountable for all claims. Contributions: Study design, analysis code, statistics and drafting by an AI agent (Claude Opus, running in Hermes Agent) under the author's direction. The author posed the research questions, proposed the lower-bound, stopping-backtest and work-versus-success framings, rated the hand-checks, edited the text and approved every claim. Errors are the author's responsibility to correct.

TL;DR: None of the inputs we tested could say reliably, before a coding-agent run started, what it would cost; most of the spread arises during the run. Two things did work: a calibrated lower bound, and, during the run, a size threshold past which runs mostly fail.


Abstract

We asked whether the cost of an AI coding-agent run can be predicted before the run starts, and if not, what can usefully be said instead. On 424,108 sessions from the public AgentLogs dataset (list-price-equivalent spend $743,263), with every pass bar committed before the data it was tested on and with repositories held out, the tested predictors mostly failed. Grouping by model, trigger and prompt length narrowed the 80% cost range in 5 of 15 groups on a locked test (26.7% of test sessions, 18.1% of spend); 4 groups came out wider than the model alone. Prompt text, repository language and size, and the user's own past runs did not narrow it further (0 of 10, 0 of 37 and 0 of 37 groups passed). An empirical 90% lower bound was calibrated on the locked test (90.0% of runs met it; per-group bounds were consistent with 90% in 10 of 15 groups) and in a secondary analysis of one external benchmark dataset (90.1%); the registered two-dataset replication could not be evaluated. During a run, the signal is strong but associational: across 171,000 runs from three public software-engineering benchmarks, runs above their group's 90th percentile of estimated size resolved their task 2.3 to 7 times less often than runs in the bottom half. Stopping runs at that threshold lowered estimated tokens per resolved task in two of three datasets (by 9% and 20%) and raised it in the third (by 3%).

1. Introduction

A weather forecast is a useful comparison. It is accurate for tomorrow and useless for next month, because small differences compound. An agent run is a sequence of model calls, each depending on the last: whether the first attempt works, whether a test fails, whether the agent gets stuck. If each step is uncertain, a forecast made before step one should be wide. This study measures how wide, and what, if anything, narrows it.

Terms.

2. Hypotheses

What started out as a single hypothesis (H1) quickly grew to multiple new theories that underwent testing.

Bars were committed to PLAN.md before the data they were tested on (commit in brackets). IDs: H = hypotheses on AgentLogs; R = replication on external data; O = outcome question on benchmark data.

IDClaimPass barRegistered
H1Groups (model x trigger x prompt length) give narrower ranges than the model alonewidth <= 0.8x model-only, 90% CI upper of that ratio < 0.8, coverage CI contains 80%; in >= 3 of 15 groupsfb27fea
H2A loop detector flags runs over-represented in the most expensive tenthratio >= 2.0 (CI lower >= 1.5) AND >= 20% of top-decile spend after onset AND hand-check false-positive rate <= 30%fb27fea
H3Flagged runs fail more oftenexploratory, no barfb27fea
H4p10 is a calibrated lower boundoverall coverage CI contains 90% AND per-group CI contains 90% in >= 2/3 of groups1ff548f
H5Prompt-text clusters narrow ranges>= 2 clusters pass the H1 width bard0dc38c; final 9f586dd
H6Repository language and size narrow ranges further>= 1/3 of groups pass8553873
H7The user's own recent runs narrow ranges further>= 1/3 of groups pass8553873
R4H1 and H4 replicate on other public datasetspass in every evaluable dataset, >= 2 evaluable2f66621
O1Tail runs resolve their task less oftentail/bottom-half resolve ratio <= 0.8 and CI upper < 0.8, in >= 2 datasets9c8a822

Analyses added after a draft review are labelled post-hoc where they appear.

3. Method

Data. risenlab/agentlogs (CC BY 4.0, revision 04013a44d3432c6654bca1dcd4a01218a9406b80): step logs of GitHub-triggered coding-agent sessions, mostly on Claude Sonnet models. All 276 log shards were parsed. Sessions with at least 4 priced calls were kept: 424,108 sessions in 30,143 repositories. The cut excludes 15,849 priced sessions with 3 or fewer calls (3.6% of priced sessions, 0.7% of priced spend) and 31,110 sessions on models without a list price. Repository metadata (1.8M repositories) joined to every kept session.

Cost. Tokens multiplied by a snapshot of OpenRouter list prices taken on 2026-10-06 (lab.py), with cached input priced separately. This is a list-price equivalent: it ignores negotiated discounts, provider routing, subscription plans and failed calls, and no public dataset was available to check it against billed amounts. Dollar figures describe the logged period at those prices, not current prices.

Splits. Development shards were explored freely; five validation shards were used once per decision; ten locked-test shards, drawn with random.Random(20261006), were used exactly once (H1, H4). Later rounds split by repository (or by task instance for benchmarks), so every tested range is scored on repositories it was not fitted on. For H7, a user's history counts only sessions created strictly before the predicted one; the same user may appear in fit and test data by design, since a deployed tool would know the user's own past.

Scoring. A group passes the range test if its width is at most 0.8x the model-only width, the 90% bootstrap CI of that ratio stays below 0.8, and the share of held-out runs inside the range is consistent with 80%. Bootstrap resamples are drawn over training repositories, so clustering within repositories is preserved; the reported upper bound is the 95th percentile (one-sided 95%, equivalently the top of a two-sided 90% interval). Coverage mainly guards against ranges that are narrow because they are wrong; width carries the claim.

Multiplicity. With 15 groups and a null effect, about 0.75 groups would be expected to pass H1's three-part bar by chance; 5 passed. H6 and H7 passed 0 of 37 each, so no correction can change those verdicts.

External data. Replication used Exgentic/agent-llm-traces-v2 (10,056 benchmark runs with reported agent cost; no licence listed) and MaxDevv/real-pi-coding-agent-traces-sessions (1,291 developer sessions; licence "other"). From these, only aggregate statistics are reported and no rows are redistributed. A third candidate, open-agent-leaderboard/traces, was dropped before any cost was read because 94% of its sessions duplicated Exgentic. The outcome question used SWE-bench/SWE-smith-trajectories (MIT), nebius/SWE-rebench-openhands-trajectories and nebius/SWE-agent-trajectories (both CC BY 4.0): 171,000 runs with a test-verified resolved label and no cost. The divisor in "characters / 4" does not affect any result: every threshold is a within-group percentile, and dividing all sizes by a constant leaves every rank unchanged.

4. Results

4.1 Cost is skewed and concentrated

Every model's mean sits well above its median (Fig. 1). The most expensive 1% of sessions carry 17% of all spend and the most expensive 10% carry 46% (Fig. 4). Within each group, sessions above the group's p90 still carry 40% of that group's spend.

Fig. 1. Cost per session by model, log scale. Box p25 to p75, line median, whiskers p10 to p90, diamond mean. 424,108 sessions.
Fig. 1. Cost per session by model, log scale. Box p25 to p75, line median, whiskers p10 to p90, diamond mean. 424,108 sessions.
Fig. 4. Share of spend against share of sessions, most expensive first. Top 1% = 17% of spend, top 10% = 46%.
Fig. 4. Share of spend against share of sessions, most expensive first. Top 1% = 17% of spend, top 10% = 46%.

4.2 Before the run: ranges

Test (registered)Result
H1, validation (7,687 sessions)6 of 15 groups pass: PASS
H1, locked test (15,349 sessions, 6,509 repositories)5 of 15 pass: PASS. The 5 cover 26.7% of test sessions and 18.1% of spend. Best group 5.4x vs 12.8x model-only. 4 groups were wider than model-only.
H5, prompt-text clusters (92,429 sessions, 13,535 repositories)0 of 10: FAIL. Two clusters met both width criteria but failed coverage on held-out repositories (73% to 78% against 80%).
H6, repository language and size (66,603 held-out sessions)0 of 37: FAIL. Median width ratio 0.99.
H7, the user's last 10 earlier runs0 of 37: FAIL. Median width ratio 0.88: directionally helpful (about 12% narrower), but no group cleared the bar with its confidence interval.
Fig. 2. Width relative to the simpler baseline, per group, on held-out data. Below 1 = narrower; red line = 0.8 bar. H1 shows point estimates from the locked test.
Fig. 2. Width relative to the simpler baseline, per group, on held-out data. Below 1 = narrower; red line = 0.8 bar. H1 shows point estimates from the locked test.

Prompt text loses even to the coarse grouping: groups have a median width of 11.3x, prompt clusters 16.4x. Many prompts are short ("fix the CI", "did you do it?") and say little about the work they will cause.

Fig. 3 shows what confidence costs. A range that contains the true cost half the time runs from 46% below to 80% above the typical cost; at 90% confidence it runs from 78% below to 309% above.

Fig. 3. Range around the group median needed for a given confidence, median over 175 groups.
Fig. 3. Range around the group median needed for a given confidence, median over 175 groups.

4.3 Before the run: the lower bound

H4 passed on the locked test (registered). Its overall bound, the p10 of all development sessions ($0.23), was met by 90.0% of the 15,349 test sessions. Per-group bounds were consistent with 90% in 10 of 15 groups.

A post-hoc comparison shows what grouping adds. Pooled over the 15 groups, the group bounds, the model-only bounds and the single global bound were all met by about 89.5% of runs. Pooled coverage is not where grouping helps. It helps per group:

For Sonnet 4.5 review requests, the global bound of $0.23 held for only 52% of runs, while the group bound ($0.07) held for 91%. Group bounds are also informative values, from $0.07 to $1.02, at 21% to 45% of the group median. Per-group detail is in revision_v02.out; the full table of 175 groups is in figures/groups_full.csv.

4.4 Why prediction stays wide

Two measurements bound what the tested predictors could achieve.

Fig. 5. Median p90/p10 within group: before the run (model x trigger) vs in hindsight (+ true call-count or token-count decile). In-sample.
Fig. 5. Median p90/p10 within group: before the run (model x trigger) vs in hindsight (+ true call-count or token-count decile). In-sample.

What this shows is narrower than impossibility. None of the inputs we tested gave reliable narrowing out of sample, and a meaningful spread remains even with information no pre-run predictor could have. A predictor using inputs we did not test, such as the repository's state or the agent's planned steps, might do better.

4.5 External validation: one evaluable dataset

The registered rule (R4) needed two independent evaluable datasets. Neither candidate met the registered minimum group size (100 fit / 30 test sessions). Under a secondary minimum of 50 / 20, added before any cost was read and labelled secondary throughout, only Exgentic qualified (60 groups); the developer-session set had one eligible group. R4 is therefore not evaluable, for both H1 and H4.

In the secondary Exgentic analysis, the lower bound held for 90.1% [89.1%, 91.4%] of held-out runs, and 48 of 60 groups were individually consistent with 90%. The range result (29 of 60 groups passing) mostly reflects model-only baselines that mix six benchmarks and are therefore very wide (up to 232x). Where a model's baseline was narrow (GPT-5.2, 6.2x), 1 of 11 groups passed and 4 were wider than baseline.

4.6 During the run: loops (H2, H3)

A detector flagged runs with long stretches of repeated identical tool calls or no file edits. On development data, flagged runs were 3.3x [3.0, 3.6] as common in the most expensive tenth as elsewhere (40% vs 12%), and 26% of top-tenth spend came after the flag. In a blind hand-check of 60 runs by one rater, 20 of 30 flagged runs were judged not to be stuck: a 67% false-positive rate against the registered 30% maximum. H2 fails. The rater's answers also drifted with case order, so the reference itself is weak.

H3 (exploratory): flagged runs ended as failed, cancelled or timed out 5.6% of the time (n = 1,206), against 5.1% for the rest (n = 7,010): essentially no difference. The flag marks expensive runs, not stuck or failing ones, and it was dropped as a loop detector.

4.7 During the run: size and failure (O1)

DatasetRunsResolve rate, tailResolve rate, bottom halfRatio [90% CI]
SWE-smith23,82120.7%51.8%0.40 [0.37, 0.43]
SWE-rebench OpenHands67,07422.6%52.9%0.43 [0.38, 0.47]
SWE-agent80,0363.5%25.4%0.14 [0.12, 0.16]

O1 passes in all three (registered). Counting turns instead of estimated size gives the same ratios. The resolve rate falls almost steadily with size decile (SWE-smith: 62% in the smallest tenth, 21% in the largest). Overall resolve rates are 44%, 48% and 16.7%. SWE-agent's lower base rate makes its ratio more extreme, but its ordering matches the other two.

This is an association. Harder tasks are both larger and more often failed, so O1 does not show that a run grows large because it is going wrong, or that stopping it would lose nothing.

Runtime alarm (post-hoc). To test the alarm as a deployed tool would run it, thresholds were set on 70% of task instances (group p90 of final estimated size) and applied, turn by turn, to the other 30%. "Similar runs" means the same model (SWE-smith, SWE-agent) or the same repository (SWE-rebench), the only grouping these datasets allow.

SWE-smithSWE-rebench OHSWE-agent
Held-out runs7,2405,50824,682
Alarm fires on9.9%10.8%10.9%
... of resolved runs (false alarms if stopping)4.5%5.3%2.4%
Resolve rate, fired vs not fired19.5% vs 44.8%19.9% vs 43.6%3.3% vs 16.7%
Share of estimated tokens still ahead when it fires (median)27%11%29%
Estimated tokens per resolved run, stopping at the alarm-9.2%+2.6%-20.3%

Whether stopping is worth it depends on the setup. In SWE-rebench the alarm fires late (11% of the run left), and stopping costs more resolved runs than it saves in tokens. A warning, which leaves the decision to a person, has no such cost; whether people act on it well is untested.

5. Discussion

What the evidence supports. A substantial part of agent-run cost was not predictable from any input we tested, and that part shows up during the run. Coarse lower bounds stay calibrated per group, and a run's size relative to similar runs is a strong in-run signal of trouble.

What can be built honestly from this:

  1. Before sending: a calibrated lower bound and a typical value per model and trigger ("at least $0.54, usually $1.29"), with the range shown as the wide thing it is.
  2. During the run: a warning when the run crosses what 90% of similar runs reached. On benchmarks, failing runs collect there. Stopping automatically is a trade whose sign depends on the setup.
  3. For long conversations: separate the work an agent does from the context it carries, and show both. All public datasets used here contain task runs; agents used as long-running conversation partners re-read their growing context on every call, a cost pattern these data cannot measure.

Limits.

Disclosures. All are logged with dates in PLAN.md.

6. Data and code

Repository: https://github.com/vhsgreed/cost-predictor-spike. Python 3.14, pyarrow 25.0.1, numpy 2.5.3, matplotlib 3.11.2, huggingface_hub 2.1.1. Inputs are public and pinned by revision; the repository contains no dataset rows. See the repository README for run order.

Change log from v0.1 (response to an LLM-generated review, 2026-10-07)