hraness
Theme
Appearance

when Ben Guo picks a frontier model

a field study of model choice at session start, and the routers it produced

by hraness · drafted with ai assistance

When Ben Guo opens a new session with a coding agent, he makes one decision before typing anything the agent will read: which model gets the task. This study reads that decision back out of his local session stores, asks what predicts it, and keeps a small classifier, a prompt router, in step with the answer. This is the second reading. The first, on September 22, 2026, found that prompt shape, not difficulty, predicted the pick. Two days later the local stores hold a different corpus and, for September, a different answer.

The corpus is every session record the three tools keep on one machine as of September 24, 2026: 22,020 Codex threads in the Codex index, of which 293 are human sessions; 1,536 Claude Code transcript files, of which 26 are top-level sessions; and 49 Devin CLI sessions. After removing harness-injected context and 16 machine-generated prompts, 337 human-authored sessions have a usable first prompt and a known first-turn model, and 91 of them fall in September, the only month in which a frontier model was available to pick. The label on each session is the model that served the first turn: claude-fable-* and gpt-6-astra count as frontier; swe-2-max, gpt-5.6-sol, gpt-5.x, claude-opus-*, and the rest of the standard tier count as standard.

labeled sessions
337
frontier share, september
51%
router held-out AUC, september
0.64
cost per routing decision
~1.1k tokens

Two findings replace the first study's. The decision has moved from the prompt to the tool: in September, 34 of 36 Codex sessions ran gpt-6-astra, 12 of 20 Claude Code sessions ran fable, and none of 35 Devin sessions left the standard tier. And within that window the pick reads as difficulty-calibrated for the first time: longer, harder, broader prompts go to the frontier, and the continuation marker that carried the first classifier is now neutral. The shipped classifier, model-router in the algal repository, was updated with a small prior-anchored step and scores 0.64 held out on the September sessions, against 0.63 for the previous head. That gain is within noise, and this page says so.

the stores rotate: what survived from the first study

The first study counted 1,251 session records and kept 321 labeled human sessions from July 30 to September 22. Reading the same stores two days later finds 125 labeled sessions in that window. Nothing was deleted by hand. Codex keeps an index row for every thread it has ever started, including the first prompt, the model, and the reasoning effort, but only 56 of the 293 human threads still have their message files; the rest expired. Claude Code keeps recent transcripts only, and 1,510 of its 1,536 files are subagent transcripts nested under 26 top-level sessions. The Devin CLI store now begins on September 10, so most of the 89 Devin sessions the first study labeled are gone, along with the per-turn model records that resolved its adaptive sessions.

Window July 30 to September 21 Codex Claude Code Devin Total
Labeled sessions, first study 195 37 89 321
Labeled sessions, today's stores 91 14 20 125

The first study's aggregates are preserved in this page's data file history and are not restated as current. Everything below is computed from today's stores, and the classifier is evaluated only on the window where the choice existed.

the corpus: what counts as a choice

Sample accounting Codex Claude Code Devin
Records in the store 22,020 threads 1,536 files 49 sessions
Human sessions with a user-authored turn 293 24 49
Machine or harness prompts excluded 11 0 5
Unlabeled (adaptive or synthetic first turn) 0 4 9
Human sessions, labeled 282 20 35
Frontier picks 34 (12%) 12 (60%) 0
September only, labeled 36 20 35
September frontier picks 34 (94%) 12 (60%) 0

The Codex index reaches back to January, and 246 of its 282 labeled sessions predate the first frontier model on this machine. Every one of them is a standard pick, which is why the whole-corpus frontier share (14%) means little: it measures when astra and fable appeared, not how they are chosen. Excluded prompts are the ones software wrote: read-only analytics probes run from /tmp, an approval-judge transcript, and empty first turns. Devin's nine adaptive sessions delegate the pick to the tool and its store no longer records which model answered, so they stay unlabeled rather than guessed.

the tool is most of the decision

Frontier share of September sessions by tool
Codex
94%
Claude Code
60%
Devin
0%
  • frontier
  • standard

In the first study Codex was a near coin flip between astra and sol. In September it is an astra shop: 34 of 36 sessions, with one gpt-5.6-sol and one codex-mini session as the whole tail. Claude Code split 12 to 8, and the 8 are deliberate: six claude-opus-5 and two claude-opus-5-5 sessions, none on sonnet. Devin stayed on swe-2-max (25) or sol variants (10). Ben's stated floor, never less than SWE-2, Sol, or Opus, holds in every September session but the one codex-mini pick.

The practical consequence for a router is that the tool chosen carries most of the information. A prompt-only classifier is asked to recover the remaining part: which Claude Code sessions go to opus, and which of the rare Codex sessions drop below astra. Weekly shares show the same thing from the other side. Frontier share was 50% in week 36 on two sessions, 100% in week 37 on five, 55% in week 38 on forty, and 41% in week 39 on forty-four: not a rising adoption curve but a settled habit with a tool mix underneath it.

what predicts the frontier pick in september

Each September prompt was scored two ways: deterministic surface features computed by the router's own program (word count, whether the first word is an action verb, continuation markers, distinct action verbs) and six typed-decision scores from TypeSafe Jev (difficulty, scope, ambiguity, stakes, task kind, and a frontier-worthiness probability). Single-feature AUCs against the observed choice, on the 91 labeled September sessions:

Signal Direction AUC
Prompt words → frontier (longer = frontier) 0.71
Jev difficulty score → frontier 0.70
Jev scope score → frontier 0.69
Jev frontier-worthiness → frontier 0.67
Distinct action verbs → frontier 0.63
Opens on an action verb → frontier 0.58
Jev ambiguity score → frontier 0.56
Jev stakes score ≈ none 0.52
Continues earlier work (resume, continue, "session named") ≈ none 0.50

Three of these reverse the first study. Length now points to frontier: the median September frontier prompt is 69 words against 27 for standard, and the frontier mean is almost twice the standard mean (115 against 59). Difficulty, which separated nothing in the first corpus (AUC 0.52), is now the second strongest signal, and the Jev means separate cleanly: difficulty 2.79 against 2.14, scope 2.83 against 2.16, frontier-worthiness 0.52 against 0.40. And the continuation marker, the single strongest feature in the first classifier, carries no information in September: 44 of the 91 sessions are resume-kind takeovers, and they split exactly in half.

Words in the first human prompt, September sessions, median by model tier
frontier
69
standard
27

Task kind tells the same story with the tool mix showing through:

Jev task kind, September Sessions Picked frontier
resume: continue or take over earlier work 44 50%
feature: build or extend 15 60%
research: investigate or report 7 71%
ops: deploy, configure, migrate 5 20%
chore: mechanical, housekeeping 4 50%
probe: capability test of the model 4 50%
bugfix 3 0%
writing: prose or editorial 3 67%
question: answer, little code 2 50%

The first study's two gates survive as policy rather than as findings. Questions were 0 of 21 frontier then and are 1 of 2 now; probes are a coin flip in both readings. Both kinds still gate to standard in the shipped program, because a capability probe does not need capability and a question rarely does.

the long prompt

The long prompt keeps its signature. Among all 337 labeled sessions, 35 first prompts run to 150 words or more (median 293, longest 752). Twenty-five of the 35 open on "i want" or "i wanna"; none ends with a question mark; like appears 3.4 times per prompt against 0.3 in prompts under 40 words, and we or let's 3.6 times against 0.2. They are prose, about 3.5 paragraphs each, with bullets creeping in at 2.7 lines per prompt against 1.2 in the first study. By Jev kind they are mostly feature briefs (14) with a few research, resume, and single-item buckets.

What changed is where they go. In the first study long prompts skewed slightly standard (32% frontier against a 36% baseline). Today 8 of 35 picked frontier (23%), but 23 of the 35 are July and August prompts with no frontier to pick. Among the 12 September long prompts, 8 picked frontier, which is the new length finding seen from the top of the distribution. The vision brief is still written the same way; it is now more likely to be sent to the frontier tier.

reasoning effort follows the model

Codex records a reasoning-effort field per thread. Every one of the 34 astra sessions ran ultra. Standard sessions spread across ultra (72), high (62), max (54), medium (34), xhigh (23), and low (3), with the lower efforts concentrated in the July and August sol sessions. Effort remains a property of the model picked, not a separate decision, so the router still needs one binary. Twenty of the 337 sessions switched models mid-session; the label is always the first turn.

Every user message that survives in the three stores was scanned for URLs: 619 link-bearing messages containing 342 unique URLs across 94 human sessions. The Codex numbers are lower-bounded, because only 56 of the 293 human threads still have their later messages; the other 237 contribute first prompts alone.

The two structural facts from the first study hold. Links are a mid-session act: 95 of the 619 link-messages arrive in a first prompt and 524 arrive once work is underway. And the largest destination is the operator's own portfolio: of 169 github.com link-messages, 111 point at hraness/* repositories, led by direct (31), hra (15), bigdatadepot (12), wrench (10), and slopcamera (8). Link-bearing sessions pick frontier at about the corpus baseline (17% against 14%), and 71 of the 94 are Codex threads, most of them from July and August.

Link destination Unique URLs Share of unique
Own infrastructure (hraness repos and product domains) 159 46%
External articles, blogs, and misc web 118 35%
GitHub, external orgs 40 12%
X posts 12 4%
Docs, packages, video 9 3%
arXiv papers 3 1%

The sibling link→project router, skills/link-router in the jungle repository, was re-evaluated on its observed session-context mappings with live typed decisions on September 24: 54 of 60 context-bearing links routed to the observed project (90%, against 90.9% on 55 in the first study). The six misses are external articles pasted in general-purpose sessions that the router filed under a plausible neighbor. No new session-context mappings were added from this corpus: the sessions whose later messages survive were mostly working in the portfolio repository itself or in a general directory, which gives no project prior worth recording.

the classifier, generation 1

The shipped router is unchanged in shape, an ALGAL organism three cells deep, examples/model-router.algal.json in the hraness/algal repo:

src (task text)
 ├─→ feats   (expr)   words, imperative, resume, verbs_c        : free
 └─→ probe   (decide) difficulty, scope, ambiguity, stakes,     : 1 typed call
                     kind, frontier-worthiness                  (~965 tok in)
        ↓
     combine  (expr)   fitted linear head + kind gates → route

The head was not refit from zero. It was updated with the rule xcb's route reflex uses to learn from an operator's own replies: full-batch logistic descent on the new examples with a Gaussian penalty of strength 24 centred on the parent weights, 600 iterations, learning rate 0.5. Eighty-four examples against a strength-24 prior move the weights a little and the intercept almost not at all:

z = -1.7683
  + 0.527 · min(words,400)/400
  + 1.403 · imperative
  + 0.449 · resume
  + 0.444 · min(distinct_verbs,8)/8
  + 0.005 · difficulty/5
  + 0.935 · scope/5
  − 0.146 · ambiguity/5
  − 0.513 · stakes/5
  + 0.611 · frontier-noul
route = frontier when z ≥ −0.619 (p ≥ 0.35)
Evaluation on the 84 labeled, non-probe September sessions Generation 0 (Sept 22 head) Generation 1 (shipped) September-only fit (not shipped)
Held-out AUC (8-fold) 0.63 0.64 0.65
Accuracy at the shipped threshold 0.58 0.58 0.62
Predicted frontier share (observed 50%) 56% 58% 81%
Weight on difficulty −0.08 +0.00 +0.72
Weight on resume +0.69 +0.45 −0.32

The last two rows carry the reading. A head fit from zero on September alone learns the new regime, difficulty up and resume down, but 84 examples do not justify replacing a 306-example parent, and its predicted share at the old threshold (81%) shows how far its calibration drifts. The anchored update moves in the same direction by a fraction of the distance. A third fit, across all 321 non-probe sessions from May to September, reaches AUC 0.82 and is not shipped either: every frontier pick is a September prompt, so that head learns the calendar through prompt-style drift rather than anything about routing.

What the router does well is unchanged: it discriminates prompt complexity with task-intrinsic features (no operator identity, repository names, or timestamps) at the cost of one bounded typed call. What it cannot do is also clearer than before: it cannot see the tool, and in September the tool is most of the decision. xcb's route reflex carries the same generation-1 head as its shipped prior and keeps learning generations from the operator's replies, which is where the next real gain will come from. The threshold remains a policy knob: raising it makes the router more conservative with frontier spend.

data, sources, and methodology

Sample construction and exclusions stores, denominators, label rule

The corpus is the three local stores on one machine as of noon UTC on September 24, 2026. Codex sessions come from the Codex thread index (state_5.sqlite, 22,020 threads), which records each thread's first user message, model, reasoning effort, and origin; threads spawned as subagents, guardian reviews, automation, or forks are excluded. Later user messages come from the surviving rollout files (51) and the thread item store (56 threads). Claude Code sessions come from the 26 top-level transcripts under ~/.claude/projects (subagent transcripts nested under them are excluded); the first-turn model is the first assistant message's model. Devin sessions come from the CLI SQLite store (49 sessions since September 10) with the configured model as the label and user messages read from the message tree.

Injected context is stripped before anything is counted (<environment_context>, # AGENTS.md instructions, system reminders, slash-command wrappers). Excluded as machine-authored: prompts run from /tmp working directories (read-only analytics probes), an approval-judge transcript, and empty first turns. Unlabeled: Devin adaptive sessions and Claude transcripts whose only assistant turns are synthetic. 337 of 350 human sessions are labeled; 327 scored through Jev; the ten unscored first prompts are all 349 words or longer and exceeded the router's decide-cell budget, and they drop out of Jev-dependent rows only.

Classifier construction features, fitting, thresholds

Features are exactly what the router's expr cell computes (word count on single-space splits, first-word verb lookup, continuation substrings, distinct action verbs capped at 8) plus the six Jev scores. Generation 1 is a maximum-a-posteriori logistic update anchored on generation 0, reproducing xcb_core::reflex::fit: full-batch gradient descent on log loss plus a Gaussian penalty of strength 24 centred on the parent, 600 iterations, learning rate 0.5, scale equal to the example count plus the prior strength. Cross-validation refits the anchored head on each of 8 shuffled folds. The September-only and full-corpus comparison fits are L2 logistic regression from zero. Probe-kind sessions are excluded from fitting and kept in reported class rates.

Long-prompt and link analyses cohorts, counting rules

The long-prompt cohort is the 35 first prompts of 150 words or more among the 337 labeled sessions; style markers are deterministic counts on the prompt text. The link corpus is every URL in every surviving user-authored message, deduplicated to 342 unique URLs. Destinations are classified structurally: github.com/hraness/* and owned product domains are own infrastructure; arxiv.org, x.com, package registries, docs hosts, and video hosts have their own buckets; everything else is external. Codex link counts are lower bounds because most older threads keep only their first prompt.

Interpretation limits what this evidence cannot show

One operator, one machine, one month in which the choice existed. The label is the model picked, not the model that would have done better; there is no counterfactual run, outcome measurement, or cost accounting. The September window holds 91 labeled sessions, so every AUC on this page carries wide uncertainty and the differences between heads are within it. Provider defaults and availability shape the picks at least as much as the prompt does. Resume-style prompts still hide their true complexity in a referenced session the classifier cannot see. Jev scores are one provider's typed judgments used as features. Store rotation means the first study's corpus cannot be reconstructed, so the two readings compare different samples of the same practice, not the same sample twice.

editorial ownership

Hraness. First study: analysis and drafting by Devin (SWE-2 Max) at Ben Guo's request, September 22 to 23, 2026. This reading: extraction, scoring, fitting, and drafting by Claude Code (Fable 5.1) at Ben Guo's request, September 24, 2026. Scoring: TypeSafe Jev through the router organism itself, about 965 input and 163 output tokens per scored prompt. The model router, its questions, weights, and replayable fixtures ship in the hraness/algal repository as examples/model-router.algal.json; xcb carries the same head as its route reflex prior; the link router ships as skills/link-router in the jungle repository. Aggregates behind this page are at /prompting/data.json. Transcripts, session identifiers, and raw prompts remain private. Reassessment due November 24, 2026.