- Conduid
- Marketplace
- #token-optimization
MCP servers tagged token-optimization
62 live MCP servers tagged "token-optimization", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with token-optimization.
Headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP s…
Laravel Toon
TOON encoding for Laravel. Encode data for AI/LLMs with ~50% fewer tokens than JSON.
Tura
Build agent that uses 80% less token and delivers better results.
Lean Ctx
Hybrid Context Optimizer — Shell Hook + MCP Server. Reduces LLM token consumption by 89-99%. Single Rust binary, zero dependencies.
LAP
Your agents are guessing at APIs. Give them the actual Agent-Native spec. 1500+ API's Ready To-Use skills, Compile any API spec into a lean, agent-native format. 10× smaller. Ope…
Claude Modular
Production-ready modular Claude Code framework with 30+ commands, token optimization, and MCP server integration. Achieves 2-10x productivity gains through systematic command or…
Agent Skills Hub
Molavi Agent Skills: Curated AI agent skills and MCP configs for Antigravity, Cursor, Codex & Claude — by Taghi Molavi
Token Saver
Content-aware output compression for AI coding assistants. 36 specialized processors cut CLI output tokens by 60-99% (git, pytest, npm, terraform, kubectl, docker, and more) witho…
Ctx
CTX - Context Runtime Engine for Coding Agents
OpenLore
openlore provides persistent architectural memory for AI coding agents by turning codebases into queryable knowledge graphs featuring static analysis, living specs, automated drif…
Omni
The Context OS for Autonomous AI Agents. Distill terminal noise into pure semantic signal, stop agent hallucinations, and cut token costs by up to 90%.
Ratel
Context engineering for AI agents. ~80% fewer tokens. Fix tool overload. Skills and memory with in-process BM25 retrieval. No vector DB. No embeddings.
Mcptoon
Token-efficient MCP CLI client. 97% less tokens on tool discovery, 40-60% on results. Zero deps. Cross-platform. Works with every AI agent.
Entroly
The Context Engineering Engine. Your AI sees 5% of your codebase — Entroly shows it everything. 78% fewer tokens. Works with Cursor, Claude Code, Copilot, OpenClaw.
Prompt Caching
Automatic prompt caching for Claude Code. Cuts token costs by up to 90% on repeated file reads, bug fix sessions, and long coding conversations - zero config.
Toonify MCP
Automatic token optimization for Claude Code and MCP workflows, including structured data and source code compression.
Stacklit
One command gives AI agents instant codebase context. ~250 tokens replaces 50,000+ tokens of exploration. Auto-configures Claude Code, Cursor, Aider.
Notebooklm Wiki Pipeline
Turn Google Drive PDFs into Obsidian wiki notes via NotebookLM MCP without loading full PDFs into Claude context
Llmtrim
Local proxy that compresses your LLM API requests so you pay less, with no change to the answers. Trims wasted tokens from prompts, history, tool output, and code before they're s…
Token Goat
Token burn reducer and focus keeper for Claude Code, Codex, Gemini CLI, Cline, Windsurf, Aider, Cursor, Copilot, and more: surgical read hints, PDF/Office/CSV/markdown file interc…
Trueline MCP
Smarter reads, safer edits. An MCP plugin that cuts token usage and catches editing mistakes before they hit disk. Supports Claude Code, Gemini CLI, GitHub Copilot, and Codex.
Claude Teams Brain
Give your Claude Code Agent Teams a memory. Auto-injects role-specific context into every new teammate — your team never starts blind again.
Reshapr
The open source, no-code MCP Server for AI-Native API Access
Token Ninja
token-ninja routes deterministic shell commands locally — zero LLM calls, ~19µs latency. Works silently inside AI tools via MCP.
Slurp
Token-budget-aware graph navigation for AI coding agents. Serve exactly the noodles your LLM needs. 🍜
PlayGuard
AI-powered MCP proxy for Playwright and Figma. Optimize Claude Code and AI agents with session recovery, context reduction and intelligent tool routing.
Terseai
The AI agent butler for macOS & Windows. Live-monitor Claude Code, Cursor, Codex, Copilot & more, stop runaway agent spend before the next API call, manage MCP servers, and cut to…
Everything Slim
everything MCP (58% less tokens). Quick setup: npx everything-slim --setup
Context Control
🔥 Context Control MCP v5.0 PORTABLE EDITION - Análisis universal de proyectos que funciona en CUALQUIER entorno. Contexto completo automático para IA.
Slim
MCP proxy that gives agents their context window back. Schema compression, lazy loading, response caching.
Larkx
AI codebase indexer and MCP server for Claude Code, Cursor, and Copilot. Pre-index your project into a compact graph and measure real token savings with `larkx bench`.
Gamedev Log Analyzer
Token-efficient game-engine & build log analysis (Unreal/Unity/Godot/MSVC-UBT-MSBuild) for the CLI and MCP — parse, dedup, classify by severity/category, diff runs, locate file:li…
Claude Code Deepseek Delegator
MCP server that lets Claude Code delegate heavy-token tasks to DeepSeek. Claude orchestrates; DeepSeek does the heavy lifting. Zero dependencies.
Distill Shrink
MCP middleware that compresses tool/resource descriptions before they hit context, with net-savings telemetry shared with the Distill skill.
Yats Toolkit
YATS — Yet Another Token Saver. One-command MCP server that indexes your codebase into a knowledge graph so AI agents (Claude, Cursor, Copilot) can search and navigate it without…
Tool Search
MCP proxy server — 85-96% token savings via lazy tool loading & fuzzy search across 100+ MCP servers. npm i mcp-tool-search
Cortex Works Minimal
A hyper-optimized MCP server that supercharges AI coding agents. Outperforms built-in IDE tools with AST-based precision editing and native OS control, drastically reducing LLM to…
Serena Slim
serena MCP (38% less tokens). Quick setup: npx serena-slim --setup
Memory Slim
memory MCP (44% less tokens). Quick setup: npx memory-slim --setup
Sequential Thinking Slim
sequential-thinking MCP (0% less tokens). Quick setup: npx sequential-thinking-slim --setup
Vibemcp
Token-Optimized Unified MCP Server for Gmail & Microsoft 365. 60% fewer tokens, 100% more power.
preflight-dev/preflight
service contract awareness, correction pattern learning, and cost estimation.
SonAIengine/graph-tool-call
tools.
Reshapr.io
reShapr website
Context7 Slim
context7 MCP (0% less tokens). Quick setup: npx context7-slim --setup
Token Scout
For OpenClaw, Hermes and more. Find free and low-cost inference (LLM models). Use them directly. Provides both a CLI and MCP server that knows which free-tier LLM APIs exist, whic…
Tokenshrink
Compress LLM prompts 30-60% — CLI, Claude Desktop MCP, browser extension. Zero API calls. Zero cost.
Alembic
Reduce noisy shell, CI, diff, and MCP-adjacent output into compact answers your coding agent can actually use. Alembic is a local, skill-first tool for Codex and Claude that cuts…
Codelens MCP Plugin
Rust MCP server for bounded code intelligence, gated mutation, and auditable agent workflows.
PsChina/deepseek-as-subagent
platform installer.