New in v0.4Xval Engine

Only what the task needs.

Xval Engine is a small model that runs on your Mac. Before Claude starts, it picks the tools and skills the task actually needs, so Claude’s context holds your work instead of listings it will never use.

The problem

A full desk before you start.

Every skill, server and tool you install is described to Claude on every turn, whether the task needs it or not. We measured how much that costs.

Context already used before the first prompt
58k tokens
Skills listed in a session, against skills invoked
129listed0invoked
MCP servers offered, against servers used
10offered1used
Input spent on listings the session never used
12.3%

Medians per session, from 30 days of our own work: 157 Claude Code sessions and 134 Codex sessions. Only counts and lengths were recorded; no conversation text was copied.

How it works

Routing, decided in milliseconds.

  1. A small model on your Mac

    A compact encoder, fine-tuned for routing and running on Core ML, reads the task. A learned scorer weighs it against what you use in this project and earlier in the session.

  2. A tiered listing

    The ten skills most likely to matter get their full descriptions. Every other skill stays listed by name, one read away, so Claude never loses sight of what you have installed.

  3. Decided at launch

    Xval routes once, when an agent starts, from its first prompt. What Claude sees then stays steady for the rest of the session.

  4. Fails open

    If Xval is missing or slow, the agent starts with its normal, full setup. Routing can make a session leaner. It never stands in the way of one.

Without Xval, every skill arrives with its full description. With Xval, the top ten keep theirs and the rest are listed by name.

How it performs

Cheaper tasks, same results.

We ran the same tasks with and without Xval, side by side, and let deterministic checkers decide whether each one was done.

15–20%

Lower cost per task

With full tool routing (in the app from v0.4.1), on hard, multi-step coding tasks, with the same success rate: 62 of 64 runs against 63 of 64 on Sonnet, and 20 of 20 against 20 of 20 on Opus.

0

Prompt-cache resets

Across 84 benchmark runs, Claude’s cached prompt was never rewritten. The saving is not given back to cache misses.

38/38

Skills found as reliably

With 141 skills installed, Sonnet completed every skill task, just as it does when every skill is listed in full, at about half the cost. On Opus, 16 of 16, at about a third less.

Cost per task, against stock Claude Code at 100
Sonnet 5.532 hard tasks, 2 runs each

15% lower

Opus 5.510 hardest tasks, 2 runs each

20% lower

Long sessionsUp to 690k tokens of context

0–3% lower

Measured on a 32-task benchmark with gold checkers, Claude Sonnet 5.5 and Opus 5.5, paired against stock Claude Code. Whiskers are 95% confidence intervals. Success is tied within what 64 runs can show; the benchmark cannot rule out a small loss, up to about 5 points. Skill results come from a separate 19-task benchmark.

Long sessions save less. Once a session grows past 600k tokens, the conversation itself is most of the bill, and Xval’s saving falls to about 0–3%. Calls to non-existent tools and recalled facts were the same with and without Xval up to 690k tokens: none invented, every fact correct. On Sonnet, Xval also missed one skill task in those sessions that stock completed.

Xval Engine ships in XAgents v0.4 as an opt-in. Turn it on in Settings › Xval Engine to see what it saves on your Mac.

The numbers

Small, fast and measured.

Replayed against the same sessions, held out by date, and timed on our own Mac.

Of the full skill listing, with every skill still visible
11%
Skill recall in the top ten, against 0.59 from usage priors alone
0.69
Recall at launch, in the top thirty, from the first prompt
0.95
The int8 Core ML model on disk
33.6 MB
To encode a prompt
5.5 ms
For a routing decision, median
10.4 ms

Privacy

It never leaves your Mac.

Xval runs entirely on your Mac. The model, what it learns from your usage and every routing decision stay there. Nothing is sent to us, or to anyone else, to decide what Claude sees.

  • The model ships inside the app and answers over a private local socket.
  • It learns from counts of what you use, not from copies of your conversations.
  • It stays off until you turn it on, and off means XAgents works exactly as before.

Coming next And one more thing it will watch

“Done” should mean done.

When an agent says it’s finished but hasn’t run the tests since its last edit, Xval notices and nudges it back to work, before you ever see a false finish.

Research paper

Coming soon

Xval: On-Device Routing of Skills and Tools for Coding Agents

The XAgents team

The full write-up: how we measured context, the router and its baselines, the replay results and their limits. We’ll link the paper here as soon as it’s published.

Leaner sessions, one switch away.

Xval Engine is in XAgents v0.4, off by default. Turn it on in Settings › Xval Engine. We’re rolling v0.4 out to the waitlist now.The XAgents team

Join the waitlist

A few short questions. Only your email is required.

Prefer email? Write to info@xagents.in.

XAgents waitlist1 of 8

Question 1 of 8: What’s your email?

We’ll use it to invite you and send occasional updates about XAgents.