1. Conduid
  2. Files
  3. DataRaum
MCP server · Files

DataRaum

Pre-computed metadata context engine for AI-driven data analytics

Unclaimed files
37Low

Scored 4 months ago · breakdown

About DataRaum

DataRaum is an MCP server in the Files category: pre-computed metadata context engine for AI-driven data analytics. It has been installed 0 times through Conduid.

Install

uvx
uvx dataraum
pip
pip install dataraum

This server has no ConduID identity, so agent calls to it are not receipted. Pin the version you install and review the source before granting it credentials.

Ask AI

Ask AI about DataRaum

Powered by Claude · Grounded in docs

I know everything about DataRaum. Ask me about installation, configuration, usage, or troubleshooting.

Security checks

  • ·README presentNot checked yet.
  • ·License declaredNot checked yet.
  • ·Tests presentNot checked yet.
  • ·Dependencies pinnedNot checked yet.
  • ·No dynamic code executionNot checked yet.
  • ·Scoped permissionsNot checked yet.

README

DataRaum Context Engine

PyPI version Python License CI

A rich metadata context engine for AI-driven data analytics.

Traditional semantic layers tell BI tools "what things are called." DataRaum tells AI "what the data means, how it behaves, how it relates, and what you can compute from it."

The core insight: AI agents don't need tools to discover metadata at runtime. They need rich, pre-computed context delivered in a format optimized for LLM consumption.

Quick Start — MCP Server

The most common way to use DataRaum is as an MCP server inside Claude Desktop (or any MCP-compatible client).

# Install
pip install dataraum

# Or with uv
uv pip install dataraum

Add to your Claude Desktop config (claude_desktop_config.json):

{
  "mcpServers": {
    "dataraum": {
      "command": "dataraum-mcp"
    }
  }
}

Then in Claude Desktop:

Add the CSV files in /path/to/my/data and measure data quality

The server runs an 18-phase analysis pipeline and makes these tools available:

Tool Description
begin_session Start an investigation session with a contract
add_source Register a data source (CSV, Parquet, JSON, or directory)
look Explore data structure, relationships, and semantic metadata
measure Measure entropy scores, readiness, and data quality
why Explain elevated entropy and propose teach suggestions
teach Extend the operation model — sole write tool (concepts, metrics, validations, ...)
query Natural language query against the data
run_sql Execute SQL directly with export support
search_snippets Discover reusable SQL patterns from prior queries and graph execution
end_session Archive workspace and end the session

Typical Workflow

add_source(name="accounting", path="/path/to/data")
  → begin_session(intent="explore data quality", contract="exploratory_analysis")
  → look()                    # Understand the data
  → measure()                 # Check quality scores and readiness
  → query("total revenue?")   # Ask questions
  → run_sql(sql="...", export_format="csv", export_name="report")
  → end_session(outcome="delivered")

Quick Start — CLI

# Run analysis pipeline (writes metadata.db + data.duckdb to ./pipeline_output)
dataraum run /path/to/data

# Inspect what was produced
dataraum dev context ./pipeline_output

See CLI Reference for all options.

What It Produces

DataRaum analyzes your data and generates:

  • Statistical metadata — distributions, cardinality, null rates, patterns
  • Semantic metadata — column roles, entity types, business terms (LLM-powered)
  • Topological metadata — relationships, join paths, hierarchies
  • Temporal metadata — granularity, gaps, seasonality, trends
  • Quality metadata — rules, scores, anomalies
  • Entropy scores — uncertainty quantification across all dimensions
  • Ontological context — domain-specific interpretation (financial, marketing, etc.)

LLM Configuration

Semantic analysis requires an Anthropic API key:

export ANTHROPIC_API_KEY="sk-..."

Configure the LLM provider in config/llm/config.yaml. See Configuration for details.

Development

git clone https://github.com/dataraum/dataraum
cd dataraum

# Install with dev dependencies (using uv)
uv sync --group dev

# Run tests
uv run pytest --testmon tests/unit -q

# Type check
uv run mypy src/

# Lint
uv run ruff check src/
uv run ruff format --check src/

Documentation

License

Apache 2.0 — see LICENSE.

README mirrored from the source repository 4 months ago. The original is authoritative.

Questions

About DataRaum

How do I install DataRaum?

Run uvx dataraum, then add the server to your MCP client's configuration. Conduid has recorded 0 installs, so the command is known to work with current clients.

Is DataRaum safe to use with an AI agent?

Its trust score is 37 out of 100 (low). Conduid hasn't run static security checks on this repository yet, so review the source yourself before granting it credentials. It has no ConduID identity yet, so agent calls to it are not receipted.

Is DataRaum still maintained?

Conduid hasn't recorded a commit date for this repository yet. Check the repository directly for recent activity.