Choose the right model with Model Right Sizer

Model Right Sizer recommends the best AI model, reasoning effort, and token budget for each stage of a task, so you get the quality you need without defaulting to the largest model for everything. It is a Claude Code plugin available from the CloudZero marketplace.

The plugin evaluates your task in two passes: a blueprint before work starts, which is a structured, schema-conformant plan that maps each stage to a recommended model tier, and a usage report after work finishes that compares actual token spend against the blueprint's predictions.

What you need

  • Claude Code v2.1 or later

Install Model Right Sizer

  1. Add the CloudZero marketplace to Claude Code. If you already added the marketplace when setting up AI Hub, skip this step.
claude plugin marketplace add cloudzero/cloudzero-claude-marketplace
  1. Install the Model Right Sizer plugin.
claude plugin install model-right-sizer@cloudzero
  1. (Optional) To have Claude consult the agent automatically on every substantive task in a specific repository, open a Claude Code session in that repository and invoke the model-right-sizer-install skill. The skill detects whether that repository uses CLAUDE.md, AGENTS.md, or both, and writes a dedicated section into whichever file(s) it finds (creating CLAUDE.md only if neither exists). It never modifies content outside that section and is safe to run more than once.

Use Model Right Sizer

Once installed, prompt the agent in natural language from any Claude Code session. The agent itself is read-only: it reads your project context and returns recommendations, but never edits or writes files.

Blueprint a task

Describe the work you plan to do. The agent breaks it into stages and recommends a model, reasoning effort, and token budget for each one. Not every stage needs the largest model; the blueprint shows you where a smaller, faster model is sufficient.

Example prompt: "Blueprint this PR: migrate the /api/billing endpoint from REST to GraphQL, add integration tests for the new schema, and update the API reference"

Keep budget honest while work is in flight

Once you start dispatching a blueprint's stages, use the budget guard skill to keep each stage's status current and flag one that's approaching its recommended token budget. It checks real spend against each stage's token_ceiling at natural checkpoints (not a fabricated live ticker), and once spend crosses 70% of the budget as of that checkpoint, surfaces the stage's own warning so you can course-correct. Checks happen at checkpoints, not continuously, so a stage that spends heavily between two checks can cross its ceiling before the next one catches it.

Example prompt: "Track budget guard status while dispatching the billing API migration blueprint's stages"

Review token spend

After completing work, ask for a usage report. The agent compares actual token spend and latency against the blueprint's predictions and flags stages that over- or under-shot, so you can refine your model choices for similar tasks in the future.

Example prompt: "Usage report on the billing API migration: we used Haiku for the tests and docs, Sonnet for the schema work"

Preview recommendations without building

To estimate token spend before committing to a task, use the dry-run skill. The agent analyzes the task and returns model recommendations for each stage without starting any work.

Example prompt: "Dry-run: add webhook support to the new billing GraphQL API"

Audit model calls already in a repo

To check model choices in code you've already shipped, rather than work you're about to plan, use the audit skill. It finds every real model call in a target repository (an SDK or API invocation, a sub-agent dispatch, an agent's model: frontmatter), decomposes each one by intent, and dry-runs each decomposed candidate independently to see whether a smaller model or tier would have covered it. It shows you the assembled findings, then opens a PR in the target repository with one schema-conformant blueprint file and a summary table, once you confirm.

This is one of two Model Right Sizer skills that write anything outside their own review step. See Read-only architecture below for the exact scope of that exception.

Example prompt: "Audit the model calls in this repo and open a PR"

Give an agent a minimal output schema

To reduce what one agent hands back to whatever calls it (an orchestrator, a skill, or a parent agent), use the schema skill. Point it at an existing agent file, or describe one you haven't written yet, and name what actually dispatches it and what that caller does with the reply. The skill prescribes the smallest typed contract that still carries everything the caller needs: a reusable shape from a small catalogue of common agent-reply patterns, typed input and output fields, and a list of what to leave out. It shows you the before-and-after size difference, then, once you confirm, writes a marked, clearly bounded schema section directly into the target agent's file.

Example prompt: "This log-triage agent just replies with a paragraph. Give it an output schema for the on-call digest skill that calls it"

How it works

Components

The plugin installs one agent and eleven companion skills. The five below cover day-to-day use. The other six are research and tuning tools for teams extending or re-calibrating the agent's own recommendations; see the plugin README for the full list.

ComponentWhat it does
model-right-sizer agentReads your project context and returns model, effort, and budget recommendations for each task stage
model-right-sizer-install skillWrites a dedicated section into a repository's CLAUDE.md, AGENTS.md, or both (whichever it finds) so the agent is consulted before and after every substantive task
model-right-sizer-dryrun skillPreviews model recommendations for a task description without building anything
model-right-sizer-audit skillFinds real model calls already shipped in a target repository, dry-runs each one, and opens a PR with a blueprint of what it found (one of two components that write outside their own review step)
model-right-sizer-schema skillPrescribes a minimal output schema for one agent's handoff to its controller, and writes it into the target agent's file once you confirm
model-right-sizer-budget-guard skillTracks a dispatched blueprint's real spend against each stage's token budget and surfaces a warning before a stage goes over

Read-only architecture

The agent's tool access is limited to reading files and fetching public model pricing. It never edits or writes files. Four of the five skills below only ever write within the narrow scope described below:

  • model-right-sizer-install writes only within the dedicated section in CLAUDE.md and/or AGENTS.md described in step 3.
  • model-right-sizer-dryrun writes nothing.
  • model-right-sizer-schema writes only within a dedicated, marker-delimited schema section in the target agent's own file (never elsewhere in that file, and never a path other than the one you named), and only after showing you the exact contract it plans to write and getting your confirmation.
  • model-right-sizer-budget-guard writes nothing to disk: it updates a routing stage's status fields in the in-flight blueprint and, once a stage crosses its warning threshold, forwards that stage's own warning to the sub-agent already dispatched for it.
  • model-right-sizer-audit is the exception: it writes one new blueprint file and opens a PR with it in the target repository, but never edits an existing file and never touches application config. It shows you the assembled blueprint before committing anything, and that confirmation step is skipped only by an explicit opt-in flag, never implicitly, no matter how the request is phrased.

Budget grounding

Each stage's token budget is derived from six rated signals (tool-call volume, content volume, cross-reference load, validation-loop iterations, context-ingestion volume, and investigative uncertainty) run through a deterministic formula, rather than estimated freehand. This keeps the reasoning behind a stage's token budget inspectable, and keeps two blueprints scored on the same signals comparable to each other.

Pricing verification

Model Right Sizer references current model pricing when making recommendations. The agent re-verifies the model lineup and prices against the model provider's pricing page each time it runs. If it cannot verify a price, it marks the figure as unverified in its output.

Direct query recommendations

Beyond model selection, the agent identifies stages where a direct data query would answer a question more reliably and at lower cost than a model call. When it detects this, it flags the stage and recommends using a query instead.

Multi-agent handoff optimization

For tasks that use multiple agents in sequence, Model Right Sizer recommends what data to pass between agents at each handoff. This keeps handoffs lean by preventing full conversation transcripts from being forwarded between stages.

ℹ️

Have questions or feedback? Reach out to your account manager.


Did this page help you?