Skip to content
verify mcp Beta VerifyMCP is currently in beta. If you notice any issues, get in touch and we’ll put it right.

tokcalc MCP Server

NPM · @TOKCALC/MCP-SERVER · 2 COMPONENTS · SCANNED OCT 2

LLM serving capacity planner: VRAM, KV cache, GPU topology, latency, and cost.

50 Trust /100
Trust breakdown (7 categories)

How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. How we score → Why this is hard to score →

Supply Chain Security48
  • Malware scan not yet available for this package.Unverified
  • No known CVEs affecting this package version or its production dependencies.Pass
  • No install/post-install scripts declared.Pass
  • 33 of 96 dependencies flagged as unhealthy. View diagnostics → Partial
Provenance & Transparency45
Schema Quality & AI Usability54
  • AI-judged instruction clarity (good).Pass
  • Context-footprint check failed: tool/resource definitions use about 2094 tokens (~190/item across 11 items; 11 tools + 0 resources), over budget; trim descriptions and params. See how to fix → Fail
  • Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management0
  • Stability check failed: the tool surface changed between 0.2.0-alpha.2 and 0.3.0: 0 tool removals, 4 breaking changes, 4 additions. See how to fix → Fail
Tool Coverage80
  • 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
  • 40% of tool parameters carry a description.Partial
Tool Safety100
  • No prompt-injection markers were found in the server instructions, tool names or descriptions we captured.Pass
  • We read all 11 captured tool definition(s), and no name or description among them implies an irreversible operation.Pass
  • An AI judge read all 12 captured unit(s) of tool text and found none that tries to manipulate the model reading it.Pass
Capabilities100
  • Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.Pass
Install

How do I install the tokcalc MCP Server server?

tokcalc MCP Server runs locally as an npm package, launched with npx -y @tokcalc/mcp-server. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.

npm · @tokcalc/mcp-server

# add to Claude Code
claude mcp add stevecrates489-commits-tokcalc -- npx -y @tokcalc/mcp-server
// .cursor/mcp.json
{
  "mcpServers": {
    "stevecrates489-commits-tokcalc": {
      "command": "npx",
      "args": [
        "-y",
        "@tokcalc/mcp-server"
      ]
    }
  }
}
// .vscode/mcp.json
{
  "servers": {
    "stevecrates489-commits-tokcalc": {
      "command": "npx",
      "args": [
        "-y",
        "@tokcalc/mcp-server"
      ]
    }
  }
}
# add to Codex CLI
codex mcp add stevecrates489-commits-tokcalc -- npx -y @tokcalc/mcp-server
// opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "stevecrates489-commits-tokcalc": {
      "type": "local",
      "command": [
        "npx",
        "-y",
        "@tokcalc/mcp-server"
      ],
      "enabled": true
    }
  }
}
# add to OpenClaw
openclaw mcp add stevecrates489-commits-tokcalc --command npx --arg -y --arg @tokcalc/mcp-server
# ~/.hermes/config.yaml
mcp_servers:
  stevecrates489-commits-tokcalc:
    command: "npx"
    args: ["-y", "@tokcalc/mcp-server"]
// ~/.netclaw/config/netclaw.json
{
  "McpServers": {
    "stevecrates489-commits-tokcalc": {
      "Transport": "stdio",
      "Command": "npx",
      "Arguments": [
        "-y",
        "@tokcalc/mcp-server"
      ]
    }
  }
}
# add to Vellum
assistant mcp add stevecrates489-commits-tokcalc -t stdio -c npx -a -y @tokcalc/mcp-server
// mcp.json
{
  "mcpServers": {
    "stevecrates489-commits-tokcalc": {
      "command": "npx",
      "args": [
        "-y",
        "@tokcalc/mcp-server"
      ]
    }
  }
}
Changelog

Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.

  • 1 Oct 26 +11
    • Known CVEs: unverified → pass ▲ security
    • Dependency health: unverified → 0.84 ▲ functional
  • 30 Sept 26 −33
    • Known CVEs: pass → unverified ▼ security
    • Malware scan: pass → unverified ▼ security
    • Stability: 0.13 → fail ▼ security
    • Dependency health: 0.84 → unverified ▼ functional
    • Schema quality: pass → fail ▼ functional
    • First check of Tool coverage: 40 functional
    • Schema quality: excellent → good functional
    • Package version: 0.2.0 → 0.3.0 functional
    • Package version: 0.2.0 → 0.2.9 functional
  • 28 Sept 26 +1
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 26 Sept 26 +6
    • Source repository: fail → pass ▲ security
    • Stability: unverified → 0.03 ▲ functional
    • Package version: 0.2.0-alpha.2 → 0.2.0 functional
  • 25 Sept 26 65

    First indexed and scored.

Diagnostics

Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.

Captured 1 Oct 2026 · Analysed npm/@tokcalc/mcp-server@0.3.0

Provenance No attestation

The registry publishes no build provenance for this version, so there is nothing to verify.

Result No attestation
Ecosystem npm

Background: How many MCP packages publish verified provenance →

Dependencies 96 packages
Packages resolved 96
Stale 33
No linked repository 1
Tree resolution Complete

Background: SBOMs and build attestations, explained →

MCP tools · 11 exposed · ~1,980 tokens

The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability. A tool's description is untrusted text the model reads on every call, which is what makes this list a security surface and not just an inventory: how tool poisoning works →

Tool Tokens
compare_gpus ~152

Compare supported GPU or cloud SKU options for the same LLM workload. Use before recommending hardware.

NameTypeReqDescription
batchSizeinteger––
gpuCountinteger––
gpusarray–Restrict comparison to specific GPU IDs (e.g. ['h100-sxm','h200-sxm']). Use list_gpus first to find IDs.
limitinteger––
modelstringyesCanonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
outputTokensinteger––
promptTokensinteger––
quantizationstring––
sortBystring––

No output schema declared.

No examples provided.

estimate_api_vs_self_host ~163

Compare monthly token-based API costs with self-hosted GPU infrastructure under stated utilization assumptions.

NameTypeReqDescription
apiInputPricenumber––
apiModelstring––
apiOutputPricenumber––
gpustringyesCanonical tokcalc GPU ID (e.g. 'h100-sxm'). Call list_gpus first if unknown.
gpuCountinteger––
inputTokensinteger––
modelstringyesCanonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
outputTokensinteger––
quantizationstring––
requestsPerDayinteger––
utilizationnumber––

No output schema declared.

No examples provided.

estimate_capacity ~402

Estimate whether an LLM-serving configuration fits in memory and can meet throughput and latency targets. Returns VRAM/KV-cache breakdown, throughput and latency ranges, concurrency, cost, confidence, assumptions, and sources.

NameTypeReqDescription
batchSizeinteger––
cacheHitRatenumber–Fraction of requests that hit the prefix cache (0–1).
cachePrefixTokensinteger–Reusable prefix length (system prompt + RAG context) that is cached across requests.
contextTokensinteger–Full context length for KV cache computation (e.g., 8192 for RAG). If omitted, KV is computed for promptTokens only. Must not exceed the model's maxContext — call list_models to check.
continuousBatchingboolean––
continuousBatchingMultipliernumber––
enginestring––
gpustringyesCanonical tokcalc GPU ID (e.g. 'h100-sxm'). Call list_gpus first if unknown.
gpuCountinteger––
kvQuantizationstring–KV cache dtype: fp16 (2 B/value), int8 (1 B) or int4 (0.5 B). Lower precision quarters KV size — decisive at 128K+ context. vLLM supports --kv-cache-dtype fp8; llama.cpp q8_0/q4_0.
modelstringyesCanonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
outputTokensinteger––
promptTokensinteger––
quantizationstring––
reasoningTokensinteger––
speculativeBoostnumber–Speculative decoding speedup multiplier (default 2.0 when useSpeculative is true).
useSpeculativeboolean–Model speculative decoding (draft model accepts ~2x tokens per step).

No output schema declared.

No examples provided.

fetch_model_spec ~159

Fetch a model's real config.json from HuggingFace and diff it against tokcalc's catalog (layers, KV heads, head_dim, vocab, max context). Catches catalog drift before it corrupts KV-cache math. Cached 24h locally; use forceRefresh to re-fetch. Works offline via cache fallback.

NameTypeReqDescription
forceRefreshboolean–Bypass the 24h cache and re-fetch from HuggingFace.
hfRepostring–HuggingFace repo override (e.g. 'Qwen/Qwen3-32B'). Guessed from the catalog when omitted.
modelIdstringyestokcalc model ID (e.g. 'qwen3-32b'). Call list_models first if unknown.

No output schema declared.

No examples provided.

find_config_for_slo ~246

INVERSE planner: given model + context + traffic volume + optional TTFT/cost/throughput ceilings, search every GPU × topology (1/2/4/8 GPUs) and return feasible configurations ranked by cost, throughput, or value. Use when the user knows their SLOs but not the hardware.

NameTypeReqDescription
batchSizeinteger–Concurrent requests the config must hold simultaneously.
contextTokensinteger–Context length the workload must serve.
maxCostPerMillionnumber–Hard ceiling on $/M output tokens. Configs above it are excluded.
maxTtftMsnumber–Hard ceiling on time-to-first-token in ms. Configs above it are excluded.
minTokensPerSecondnumber–Minimum aggregate tok/s the config must deliver at the given batchSize.
modelstringyesCanonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
quantizationstring––
requestsPerDayinteger–Expected daily request volume (drives the monthly cost estimate).
sortBystring––

No output schema declared.

No examples provided.

get_mlperf_benchmarks ~126

Retrieve curated MLPerf Inference v4.1 LLM benchmark reference configurations. Throughput is computed by tokcalc formulas (confidence_tier=derived). Use to cross-validate theoretical estimates against audited system configurations.

NameTypeReqDescription
gpuModelstring–Filter by GPU model substring (e.g. 'H100', 'H200', 'A100')
scenariostring–Filter by MLPerf scenario
workloadstring–Filter by model ID substring (e.g. 'llama3-70b', 'llama3-8b')

No output schema declared.

No examples provided.

list_gpus ~62

List GPU and cloud SKU IDs, memory, bandwidth, pricing, and categories.

NameTypeReqDescription
categorystring––
minVramGbnumber––
vendorstring–Filter by vendor (e.g. 'NVIDIA', 'AMD')

No output schema declared.

No examples provided.

list_models ~67

List tokcalc-supported model IDs and metadata. Use before estimating if the requested model is ambiguous or unknown.

NameTypeReqDescription
categorystring––
familystring–Filter by model family (e.g. 'Llama', 'Qwen')
isMoEboolean––

No output schema declared.

No examples provided.

plan_deployment ~241

One-call deployment decision brief for a specific config: memory breakdown with KV headroom, performance, rig cost, blended build-vs-buy with break-even, concrete risks (no-NVLink TP, tight headroom, long context), and next steps. Chains what would otherwise take 4 separate tool calls.

NameTypeReqDescription
apiInputPricenumber––
apiModelstring––
apiOutputPricenumber––
batchSizeinteger––
continuousBatchingboolean––
continuousBatchingMultipliernumber––
enginestring––
gpustringyesCanonical tokcalc GPU ID (e.g. 'h100-sxm'). Call list_gpus first if unknown.
gpuCountinteger––
modelstringyesCanonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
outputTokensinteger––
promptTokensinteger––
quantizationstring––
reasoningTokensinteger––
requestsPerDayinteger––

No output schema declared.

No examples provided.

recommend_topology ~107

Recommend feasible GPU/topology designs including tensor parallelism and context parallel (RingAttention) when needed.

NameTypeReqDescription
batchSizeinteger––
contextTokensinteger–Context length to size the KV cache for. Must not exceed the model's maxContext — call list_models to check.
modelstringyesCanonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
quantizationstring––

No output schema declared.

No examples provided.

record_measured ~255

Record a real-world measurement (observed tok/s, TTFT, ITL) for a model+GPU+quantization cell. Compares against the model's prediction, stores a measured/predicted ratio, and future estimate_capacity / find_config_for_slo results for that cell are annotated with your calibration. The estimator improves as you use it.

NameTypeReqDescription
batchSizeinteger––
enginestring––
gpustringyesCanonical tokcalc GPU ID (e.g. 'h100-sxm'). Call list_gpus first if unknown.
gpuCountinteger––
modelstringyesCanonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
notesstring––
observedobjectyesMeasured values from your serving run (vLLM logs, benchmark harness, etc.). At least one field.
outputTokensinteger––
promptTokensinteger––
quantizationstring––
sourcestring–Where the measurement came from (e.g. 'vllm bench_serving, 2026-09-28').

No output schema declared.

No examples provided.

Common questions

What is the tokcalc MCP Server server?

tokcalc MCP Server is listed in the public MCP registry as io.github.stevecrates489-commits/tokcalc. LLM serving capacity planner: VRAM, KV cache, GPU topology, latency, and cost. This page covers its npm package (@tokcalc/mcp-server).

Is the tokcalc MCP Server server safe to use?

tokcalc MCP Server scores 50 out of 100 on VerifyMCP. We found no known CVEs affecting it as of 1 October 2026. It declares no install or post-install scripts. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.

What tools does the tokcalc MCP Server server expose?

tokcalc MCP Server exposes 11 tools: estimate_capacity, compare_gpus, recommend_topology, estimate_api_vs_self_host, list_models, and 6 more. Their descriptions and schemas cost roughly 1,980 tokens of context every time the server is loaded.

Is the tokcalc MCP Server server still maintained?

tokcalc MCP Server is still listed as active in the MCP registry. We last reached this channel on 1 October 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.

What licence is the tokcalc MCP Server server under?

tokcalc MCP Server declares the Apache-2.0 licence, which is OSI-approved. That covers the source only, and says nothing about the cost of any service it calls.