Internal dogfood build — illustrative, not vendor-certified

Skills are the compounding variable in coding-agent performance.

SkillLift Lab is BenchTag's running instrument for one question: when a coding agent can retrieve and use a matched skill, how much does that actually move the score — and how sure are we? Every number on this page is reported as a range, not a point.

Run window: 2026-06 → 2026-08 · 3 agents · 4 task categories · 192 tasks · updated 2026-08-09

+15.0pts
median lift, with skills vs. without — live for the current scenario
1,156
task runs logged this window
3
agents benchmarked
4
categories: frontend, data, science, ops
§ 01 Scenario

Turn the two dials the benchmark actually varies.

Every range in the leaderboard and category charts below is recomputed live from these two settings — nothing here is a static screenshot. Hover or focus any number once you've moved a slider to see exactly what moved and why.

Established · 50%
ThinEstablishedDeep

How much of the promoted skill library actually matches this task mix. Moves the with-skills side only — there's nothing to retrieve in the no-skills baseline, so that column never reacts to this slider.

Standard · 50%
EasierStandardHarder

How hard this run's task mix is relative to the reference suite. Moves both conditions and widens both ranges — but hits the no-skills baseline harder and faster than the skilled one.

§ 02 The model

Agent, Environment, Skills — three dials, one score.

Click a circle, or the center. The score you see anywhere else on this page is what happens when these three stop being independent.

Compounding zone

Agent, environment, and skills are not additive — they multiply. A strong agent with full tool access and a matched skill clears tasks that any single factor, maxed out alone, cannot. That compounding zone is what SkillLift Lab measures: not "how good is the model," but "how much of its ceiling did it actually reach."

+15.0 ptsmedian lift attributed to the compounding zone, all agents pooled
§ 03 Leaderboard coverage 50% · difficulty 50%

Score ranges, with skills loaded vs. without.

Each range is the interquartile spread across logged runs — not a mean, not a best-of. Sort by any column; toggle agents off to isolate one comparison.

Agents
SkillLift Lab · overall score range by agent · n = task runs this window
without skills (IQR) with skills (IQR) Score axis: 40 – 95, shared across every chart on this page
§ 04 Pipeline

Where "with skills" scores come from.

Skills in this benchmark aren't hand-authored once and frozen. Agents draft them from their own transcripts, and a skill only survives if it clears a held-out validation bar.

6 stages · click one, or run the loop
§ 05 Categories coverage 50% · difficulty 50%

Lift isn't uniform across domains.

Same three agents, same skill pipeline, four task categories. Where a category's skill library is deep, the gap widens; where it's thin, baseline and skilled runs sit closer together.

Agents
SkillLift Lab · score range by category × agent
Category Agent Without skills With skills Lift
§ 06 Methodology & uncertainty

Read the ranges before you read the ranking.

This section exists because a single leaderboard number invites more confidence than the underlying data supports. Here's exactly what's behind each one.

Task suite

48 tasks per category (192 total), drawn from real repository issues across frontend, data, science, and ops work. Each task ships with a rubric, not a pass/fail flag.

Scoring

0–100 rubric, graded automatically where the task allows it (tests passing, lint clean, output matching a spec), spot-checked by a second reviewer pass on a 15% sample.

Run protocol

Every agent × condition × task ran multiple times at default sampling settings — no cherry-picked seeds, and failed runs are counted, not dropped.

What "range" means

The band reported is the 25th–75th percentile across runs, not a mean and not a single best attempt. A narrow band means the score was reproducible; a wide one means the task was sensitive to how the run happened to go.

How to read "78–84"

It does not mean "this agent scores 81." It means: in the middle half of logged runs, the score landed somewhere in that band. The other half landed outside it — in either direction. Treat the midpoint as a rough summary for sorting, never as a promised score on your task.

Where two agents' bands overlap, don't read the ranking as settled — read it as "not yet distinguishable at this sample size."

Known limitations

This is a dogfood instrument, not a certification. SkillLift Lab is BenchTag's internal tool for watching our own skill pipeline. It has not been reviewed by, or produced in partnership with, Anthropic, OpenAI, or Google — the agent names above are shown only to identify which coding agent produced each run. Scores reflect this harness, this task suite, and this run window; they may not reproduce elsewhere.