Skills are the compounding variable in coding-agent performance.
SkillLift Lab is BenchTag's running instrument for one question: when a coding agent can retrieve and use a matched skill, how much does that actually move the score — and how sure are we? Every number on this page is reported as a range, not a point.
Turn the two dials the benchmark actually varies.
Every range in the leaderboard and category charts below is recomputed live from these two settings — nothing here is a static screenshot. Hover or focus any number once you've moved a slider to see exactly what moved and why.
How much of the promoted skill library actually matches this task mix. Moves the with-skills side only — there's nothing to retrieve in the no-skills baseline, so that column never reacts to this slider.
How hard this run's task mix is relative to the reference suite. Moves both conditions and widens both ranges — but hits the no-skills baseline harder and faster than the skilled one.
Agent, Environment, Skills — three dials, one score.
Click a circle, or the center. The score you see anywhere else on this page is what happens when these three stop being independent.
Agent, environment, and skills are not additive — they multiply. A strong agent with full tool access and a matched skill clears tasks that any single factor, maxed out alone, cannot. That compounding zone is what SkillLift Lab measures: not "how good is the model," but "how much of its ceiling did it actually reach."
Score ranges, with skills loaded vs. without.
Each range is the interquartile spread across logged runs — not a mean, not a best-of. Sort by any column; toggle agents off to isolate one comparison.
Where "with skills" scores come from.
Skills in this benchmark aren't hand-authored once and frozen. Agents draft them from their own transcripts, and a skill only survives if it clears a held-out validation bar.
Lift isn't uniform across domains.
Same three agents, same skill pipeline, four task categories. Where a category's skill library is deep, the gap widens; where it's thin, baseline and skilled runs sit closer together.
| Category | Agent | Without skills | With skills | Lift |
|---|
Read the ranges before you read the ranking.
This section exists because a single leaderboard number invites more confidence than the underlying data supports. Here's exactly what's behind each one.
Task suite
48 tasks per category (192 total), drawn from real repository issues across frontend, data, science, and ops work. Each task ships with a rubric, not a pass/fail flag.
Scoring
0–100 rubric, graded automatically where the task allows it (tests passing, lint clean, output matching a spec), spot-checked by a second reviewer pass on a 15% sample.
Run protocol
Every agent × condition × task ran multiple times at default sampling settings — no cherry-picked seeds, and failed runs are counted, not dropped.
What "range" means
The band reported is the 25th–75th percentile across runs, not a mean and not a single best attempt. A narrow band means the score was reproducible; a wide one means the task was sensitive to how the run happened to go.
How to read "78–84"
It does not mean "this agent scores 81." It means: in the middle half of logged runs, the score landed somewhere in that band. The other half landed outside it — in either direction. Treat the midpoint as a rough summary for sorting, never as a promised score on your task.
Where two agents' bands overlap, don't read the ranking as settled — read it as "not yet distinguishable at this sample size."