15–20%
Lower cost per task
With full tool routing (in the app from v0.4.1), on hard, multi-step coding tasks, with the same success rate: 62 of 64 runs against 63 of 64 on Sonnet, and 20 of 20 against 20 of 20 on Opus.
New in v0.4Xval Engine
Xval Engine is a small model that runs on your Mac. Before Claude starts, it picks the tools and skills the task actually needs, so Claude’s context holds your work instead of listings it will never use.
The problem
Every skill, server and tool you install is described to Claude on every turn, whether the task needs it or not. We measured how much that costs.
Medians per session, from 30 days of our own work: 157 Claude Code sessions and 134 Codex sessions. Only counts and lengths were recorded; no conversation text was copied.
How it works
A compact encoder, fine-tuned for routing and running on Core ML, reads the task. A learned scorer weighs it against what you use in this project and earlier in the session.
The ten skills most likely to matter get their full descriptions. Every other skill stays listed by name, one read away, so Claude never loses sight of what you have installed.
Xval routes once, when an agent starts, from its first prompt. What Claude sees then stays steady for the rest of the session.
If Xval is missing or slow, the agent starts with its normal, full setup. Routing can make a session leaner. It never stands in the way of one.
How it performs
We ran the same tasks with and without Xval, side by side, and let deterministic checkers decide whether each one was done.
15–20%
With full tool routing (in the app from v0.4.1), on hard, multi-step coding tasks, with the same success rate: 62 of 64 runs against 63 of 64 on Sonnet, and 20 of 20 against 20 of 20 on Opus.
0
Across 84 benchmark runs, Claude’s cached prompt was never rewritten. The saving is not given back to cache misses.
38/38
With 141 skills installed, Sonnet completed every skill task, just as it does when every skill is listed in full, at about half the cost. On Opus, 16 of 16, at about a third less.
15% lower
20% lower
0–3% lower
Measured on a 32-task benchmark with gold checkers, Claude Sonnet 5.5 and Opus 5.5, paired against stock Claude Code. Whiskers are 95% confidence intervals. Success is tied within what 64 runs can show; the benchmark cannot rule out a small loss, up to about 5 points. Skill results come from a separate 19-task benchmark.
Long sessions save less. Once a session grows past 600k tokens, the conversation itself is most of the bill, and Xval’s saving falls to about 0–3%. Calls to non-existent tools and recalled facts were the same with and without Xval up to 690k tokens: none invented, every fact correct. On Sonnet, Xval also missed one skill task in those sessions that stock completed.
Xval Engine ships in XAgents v0.4 as an opt-in. Turn it on in Settings › Xval Engine to see what it saves on your Mac.
The numbers
Replayed against the same sessions, held out by date, and timed on our own Mac.
Privacy
Xval runs entirely on your Mac. The model, what it learns from your usage and every routing decision stay there. Nothing is sent to us, or to anyone else, to decide what Claude sees.
Coming next And one more thing it will watch
When an agent says it’s finished but hasn’t run the tests since its last edit, Xval notices and nudges it back to work, before you ever see a false finish.
The full write-up: how we measured context, the router and its baselines, the replay results and their limits. We’ll link the paper here as soon as it’s published.
Xval Engine is in XAgents v0.4, off by default. Turn it on in Settings › Xval Engine. We’re rolling v0.4 out to the waitlist now.The XAgents team
A few short questions. Only your email is required.
Prefer email? Write to info@xagents.in.