# tokcalc MCP Server (npm · @tokcalc/mcp-server)

LLM serving capacity planner: VRAM, KV cache, GPU topology, latency, and cost.

- Trust score: 50/100 (low)
- Registry status: active
- Liveness: live
- Owner verified: no
- Last scored: 2026-10-01

## Components

- remote · `tokcalc.vercel.app`: 25/100, [markdown](https://verifymcp.io/servers/stevecrates489-commits-tokcalc/api-mcp.md), [page](https://verifymcp.io/servers/stevecrates489-commits-tokcalc/api-mcp)
- npm · `@tokcalc/mcp-server`: 50/100 (this document), [markdown](https://verifymcp.io/servers/stevecrates489-commits-tokcalc/tokcalc-mcp-server.md), [page](https://verifymcp.io/servers/stevecrates489-commits-tokcalc/tokcalc-mcp-server)

## Channel facts

- Registry: `npm`
- Package: `@tokcalc/mcp-server`
- Version: `0.3.0`
- Transport: `stdio`

## Trust breakdown

How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. Scores are 0–100 per category. Scoring method: https://verifymcp.io/docs/scoring (what has changed: https://verifymcp.io/docs/scoring/changelog)

Scored 2026-10-01.

- **Supply Chain Security**: 48/100
  - Malware scan not yet available for this package.
  - No known CVEs affecting this package version or its production dependencies.
  - No install/post-install scripts declared.
  - 33 of 96 dependencies flagged as unhealthy.
- **Provenance & Transparency**: 45/100
  - Source repository is publicly reachable at the declared URL.
  - Provenance check failed: no build-provenance attestation is published.
  - Clear OSI-approved license (Apache-2.0).
  - Actively maintained (last published 0 days ago).
  - Disclosure check failed: no security disclosure policy was found in the source repository.
- **Schema Quality & AI Usability**: 54/100
  - AI-judged instruction clarity (good).
  - Context-footprint check failed: tool/resource definitions use about 2094 tokens (~190/item across 11 items; 11 tools + 0 resources), over budget; trim descriptions and params.
  - Usage-examples check failed: none of the tools include examples.
- **Stability & Change Management**: 0/100
  - Stability check failed: the tool surface changed between 0.2.0-alpha.2 and 0.3.0: 0 tool removals, 4 breaking changes, 4 additions.
- **Tool Coverage**: 80/100
  - 100% of tools have a non-trivial description (not blank, and not just the tool's name).
  - 40% of tool parameters carry a description.
- **Tool Safety**: 100/100
  - No prompt-injection markers were found in the server instructions, tool names or descriptions we captured.
  - We read all 11 captured tool definition(s), and no name or description among them implies an irreversible operation.
  - An AI judge read all 12 captured unit(s) of tool text and found none that tries to manipulate the model reading it.
- **Capabilities**: 100/100
  - Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.

## Install

### How do I install the tokcalc MCP Server server?

tokcalc MCP Server runs locally as an npm package, launched with npx -y @tokcalc/mcp-server. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.

### Claude

```bash
claude mcp add stevecrates489-commits-tokcalc -- npx -y @tokcalc/mcp-server
```

### Cursor

```json
{
  "mcpServers": {
    "stevecrates489-commits-tokcalc": {
      "command": "npx",
      "args": [
        "-y",
        "@tokcalc/mcp-server"
      ]
    }
  }
}
```

### VS Code

```json
{
  "servers": {
    "stevecrates489-commits-tokcalc": {
      "command": "npx",
      "args": [
        "-y",
        "@tokcalc/mcp-server"
      ]
    }
  }
}
```

### Codex

```bash
codex mcp add stevecrates489-commits-tokcalc -- npx -y @tokcalc/mcp-server
```

### opencode

```json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "stevecrates489-commits-tokcalc": {
      "type": "local",
      "command": [
        "npx",
        "-y",
        "@tokcalc/mcp-server"
      ],
      "enabled": true
    }
  }
}
```

### OpenClaw

```bash
openclaw mcp add stevecrates489-commits-tokcalc --command npx --arg -y --arg @tokcalc/mcp-server
```

### Hermes

```yaml
mcp_servers:
  stevecrates489-commits-tokcalc:
    command: "npx"
    args: ["-y", "@tokcalc/mcp-server"]
```

### Netclaw

```json
{
  "McpServers": {
    "stevecrates489-commits-tokcalc": {
      "Transport": "stdio",
      "Command": "npx",
      "Arguments": [
        "-y",
        "@tokcalc/mcp-server"
      ]
    }
  }
}
```

### Vellum

```bash
assistant mcp add stevecrates489-commits-tokcalc -t stdio -c npx -a -y @tokcalc/mcp-server
```

### Other

```json
{
  "mcpServers": {
    "stevecrates489-commits-tokcalc": {
      "command": "npx",
      "args": [
        "-y",
        "@tokcalc/mcp-server"
      ]
    }
  }
}
```

## Changelog

Every change recorded for this component, newest first. Days that predate change tracking, or that we cannot explain, say so: "we were watching and nothing happened" and "we were not watching" are different claims.

### 2026-10-01 (score 50, +11)

- [security improvement] Known CVEs: unverified → pass
- [functional improvement] Dependency health: unverified → 0.84

### 2026-09-30 (score 39, −33)

- [security regression] Known CVEs: pass → unverified
- [security regression] Malware scan: pass → unverified
- [security regression] Stability: 0.13 → fail
- [functional regression] Dependency health: 0.84 → unverified
- [functional regression] Schema quality: pass → fail
- [functional] First check of Tool coverage: 40
- [functional] Schema quality: excellent → good
- [functional] Package version: 0.2.0 → 0.3.0
- [functional] Package version: 0.2.0 → 0.2.9

### 2026-09-28 (score 72, +1)

- [functional] We updated how we score, so this day's move reflects our rubric, not a change to the server

### 2026-09-26 (score 71, +6)

- [security improvement] Source repository: fail → pass
- [functional improvement] Stability: unverified → 0.03
- [functional] Package version: 0.2.0-alpha.2 → 0.2.0

### 2026-09-25 (score 65)

First indexed and scored.

## MCP tools (11)

### `estimate_capacity` (~402 tokens)

Estimate whether an LLM-serving configuration fits in memory and can meet throughput and latency targets. Returns VRAM/KV-cache breakdown, throughput and latency ranges, concurrency, cost, confidence, assumptions, and sources.

Input parameters:

- `batchSize` (integer)
- `cacheHitRate` (number): Fraction of requests that hit the prefix cache (0–1).
- `cachePrefixTokens` (integer): Reusable prefix length (system prompt + RAG context) that is cached across requests.
- `contextTokens` (integer): Full context length for KV cache computation (e.g., 8192 for RAG). If omitted, KV is computed for promptTokens only. Must not exceed the model's maxContext — call list_models to check.
- `continuousBatching` (boolean)
- `continuousBatchingMultiplier` (number)
- `engine` (string)
- `gpu` (string, required): Canonical tokcalc GPU ID (e.g. 'h100-sxm'). Call list_gpus first if unknown.
- `gpuCount` (integer)
- `kvQuantization` (string): KV cache dtype: fp16 (2 B/value), int8 (1 B) or int4 (0.5 B). Lower precision quarters KV size — decisive at 128K+ context. vLLM supports --kv-cache-dtype fp8; llama.cpp q8_0/q4_0.
- `model` (string, required): Canonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
- `outputTokens` (integer)
- `promptTokens` (integer)
- `quantization` (string)
- `reasoningTokens` (integer)
- `speculativeBoost` (number): Speculative decoding speedup multiplier (default 2.0 when useSpeculative is true).
- `useSpeculative` (boolean): Model speculative decoding (draft model accepts ~2x tokens per step).

### `compare_gpus` (~152 tokens)

Compare supported GPU or cloud SKU options for the same LLM workload. Use before recommending hardware.

Input parameters:

- `batchSize` (integer)
- `gpuCount` (integer)
- `gpus` (array): Restrict comparison to specific GPU IDs (e.g. ['h100-sxm','h200-sxm']). Use list_gpus first to find IDs.
- `limit` (integer)
- `model` (string, required): Canonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
- `outputTokens` (integer)
- `promptTokens` (integer)
- `quantization` (string)
- `sortBy` (string)

### `recommend_topology` (~107 tokens)

Recommend feasible GPU/topology designs including tensor parallelism and context parallel (RingAttention) when needed.

Input parameters:

- `batchSize` (integer)
- `contextTokens` (integer): Context length to size the KV cache for. Must not exceed the model's maxContext — call list_models to check.
- `model` (string, required): Canonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
- `quantization` (string)

### `estimate_api_vs_self_host` (~163 tokens)

Compare monthly token-based API costs with self-hosted GPU infrastructure under stated utilization assumptions.

Input parameters:

- `apiInputPrice` (number)
- `apiModel` (string)
- `apiOutputPrice` (number)
- `gpu` (string, required): Canonical tokcalc GPU ID (e.g. 'h100-sxm'). Call list_gpus first if unknown.
- `gpuCount` (integer)
- `inputTokens` (integer)
- `model` (string, required): Canonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
- `outputTokens` (integer)
- `quantization` (string)
- `requestsPerDay` (integer)
- `utilization` (number)

### `list_models` (~67 tokens)

List tokcalc-supported model IDs and metadata. Use before estimating if the requested model is ambiguous or unknown.

Input parameters:

- `category` (string)
- `family` (string): Filter by model family (e.g. 'Llama', 'Qwen')
- `isMoE` (boolean)

### `list_gpus` (~62 tokens)

List GPU and cloud SKU IDs, memory, bandwidth, pricing, and categories.

Input parameters:

- `category` (string)
- `minVramGb` (number)
- `vendor` (string): Filter by vendor (e.g. 'NVIDIA', 'AMD')

### `get_mlperf_benchmarks` (~126 tokens)

Retrieve curated MLPerf Inference v4.1 LLM benchmark reference configurations. Throughput is computed by tokcalc formulas (confidence_tier=derived). Use to cross-validate theoretical estimates against audited system configurations.

Input parameters:

- `gpuModel` (string): Filter by GPU model substring (e.g. 'H100', 'H200', 'A100')
- `scenario` (string): Filter by MLPerf scenario
- `workload` (string): Filter by model ID substring (e.g. 'llama3-70b', 'llama3-8b')

### `find_config_for_slo` (~246 tokens)

INVERSE planner: given model + context + traffic volume + optional TTFT/cost/throughput ceilings, search every GPU × topology (1/2/4/8 GPUs) and return feasible configurations ranked by cost, throughput, or value. Use when the user knows their SLOs but not the hardware.

Input parameters:

- `batchSize` (integer): Concurrent requests the config must hold simultaneously.
- `contextTokens` (integer): Context length the workload must serve.
- `maxCostPerMillion` (number): Hard ceiling on $/M output tokens. Configs above it are excluded.
- `maxTtftMs` (number): Hard ceiling on time-to-first-token in ms. Configs above it are excluded.
- `minTokensPerSecond` (number): Minimum aggregate tok/s the config must deliver at the given batchSize.
- `model` (string, required): Canonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
- `quantization` (string)
- `requestsPerDay` (integer): Expected daily request volume (drives the monthly cost estimate).
- `sortBy` (string)

### `plan_deployment` (~241 tokens)

One-call deployment decision brief for a specific config: memory breakdown with KV headroom, performance, rig cost, blended build-vs-buy with break-even, concrete risks (no-NVLink TP, tight headroom, long context), and next steps. Chains what would otherwise take 4 separate tool calls.

Input parameters:

- `apiInputPrice` (number)
- `apiModel` (string)
- `apiOutputPrice` (number)
- `batchSize` (integer)
- `continuousBatching` (boolean)
- `continuousBatchingMultiplier` (number)
- `engine` (string)
- `gpu` (string, required): Canonical tokcalc GPU ID (e.g. 'h100-sxm'). Call list_gpus first if unknown.
- `gpuCount` (integer)
- `model` (string, required): Canonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
- `outputTokens` (integer)
- `promptTokens` (integer)
- `quantization` (string)
- `reasoningTokens` (integer)
- `requestsPerDay` (integer)

### `fetch_model_spec` (~159 tokens)

Fetch a model's real config.json from HuggingFace and diff it against tokcalc's catalog (layers, KV heads, head_dim, vocab, max context). Catches catalog drift before it corrupts KV-cache math. Cached 24h locally; use forceRefresh to re-fetch. Works offline via cache fallback.

Input parameters:

- `forceRefresh` (boolean): Bypass the 24h cache and re-fetch from HuggingFace.
- `hfRepo` (string): HuggingFace repo override (e.g. 'Qwen/Qwen3-32B'). Guessed from the catalog when omitted.
- `modelId` (string, required): tokcalc model ID (e.g. 'qwen3-32b'). Call list_models first if unknown.

### `record_measured` (~255 tokens)

Record a real-world measurement (observed tok/s, TTFT, ITL) for a model+GPU+quantization cell. Compares against the model's prediction, stores a measured/predicted ratio, and future estimate_capacity / find_config_for_slo results for that cell are annotated with your calibration. The estimator improves as you use it.

Input parameters:

- `batchSize` (integer)
- `engine` (string)
- `gpu` (string, required): Canonical tokcalc GPU ID (e.g. 'h100-sxm'). Call list_gpus first if unknown.
- `gpuCount` (integer)
- `model` (string, required): Canonical tokcalc model ID (e.g. 'llama3-8b'). Call list_models first if unknown.
- `notes` (string)
- `observed` (object, required): Measured values from your serving run (vLLM logs, benchmark harness, etc.). At least one field.
- `outputTokens` (integer)
- `promptTokens` (integer)
- `quantization` (string)
- `source` (string): Where the measurement came from (e.g. 'vllm bench_serving, 2026-09-28').

## Diagnostics

Captured diagnostic sections: Provenance, Dependencies. The full working is on the page: https://verifymcp.io/servers/stevecrates489-commits-tokcalc/tokcalc-mcp-server#diagnostics

## Score history

- 2026-10-01: 50
- 2026-09-30: 39
- 2026-09-29: 72
- 2026-09-28: 72
- 2026-09-27: 71
- 2026-09-26: 71
- 2026-09-25: 65

## Common questions

### What is the tokcalc MCP Server server?

tokcalc MCP Server is listed in the public MCP registry as io.github.stevecrates489-commits/tokcalc. LLM serving capacity planner: VRAM, KV cache, GPU topology, latency, and cost. This page covers its npm package (@tokcalc/mcp-server).

### Is the tokcalc MCP Server server safe to use?

tokcalc MCP Server scores 50 out of 100 on VerifyMCP. We found no known CVEs affecting it as of 1 October 2026. It declares no install or post-install scripts. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.

### What tools does the tokcalc MCP Server server expose?

tokcalc MCP Server exposes 11 tools: estimate_capacity, compare_gpus, recommend_topology, estimate_api_vs_self_host, list_models, and 6 more. Their descriptions and schemas cost roughly 1,980 tokens of context every time the server is loaded.

### Is the tokcalc MCP Server server still maintained?

tokcalc MCP Server is still listed as active in the MCP registry. We last reached this channel on 1 October 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.

### What licence is the tokcalc MCP Server server under?

tokcalc MCP Server declares the Apache-2.0 licence, which is OSI-approved. That covers the source only, and says nothing about the cost of any service it calls.

## Links

- npm package: https://www.npmjs.com/package/@tokcalc/mcp-server
- Socket report: https://socket.dev/npm/package/@tokcalc/mcp-server
- Repository: https://github.com/stevecrates489-commits/tokcalc
- Changelog RSS feed: https://verifymcp.io/servers/stevecrates489-commits-tokcalc/tokcalc-mcp-server.xml
- Changelog JSON feed: https://verifymcp.io/servers/stevecrates489-commits-tokcalc/tokcalc-mcp-server.json
- HTML version of this page: https://verifymcp.io/servers/stevecrates489-commits-tokcalc/tokcalc-mcp-server
