- Conduid
- Marketplace
- #gguf
MCP servers tagged gguf
10 live MCP servers tagged "gguf", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with gguf.
Comfyui LLM Party
LLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt nodes,access to Feishu,discord,and adapts to all llms with similar openai…
OGAM
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and loc…
Atomic Chat
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.
Atomic Agent
Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.
Canary
Deterministic, read-only static auditor for GGUF models: behavioral chat-template backdoors + SSTI, tokenizer, config, and model-card surfaces - never renders, never loads weights…
Unch
Local-first semantic code search for repository annotations via GGUF embeddings and sqlite-vec.
Lilbee
Run local AI models, search your files and code, and crawl the web, all in one program. Cited answers, local-first, with an MCP server for your coding agent. TUI, CLI, REST API, a…
Compress Tokens
MCP server that compresses text by removing unnecessary tokens using local LLM surprisal scoring
ShipItAndPray/mcp-turboquant
LLM quantization via tool call. Convert models to GGUF, GPTQ, and AWQ formats. Recommend optimal quant settings, evaluate quality, and push to Hugging Face Hub.
Diffusion LLM MCP
FastMCP fleet MCP server for diffusion LMs (dLLM). DiffusionGemma on Goliath RTX 4090 — batch inference, HLE-shaped reasoning, ~200–400 tok/s. Doc phase; llama-diffusion-cli sidec…