# Pyspark Mcp

> SQL to PySpark conversion, AWS Glue job generation, and Spark code optimization.

- **Type:** MCP server
- **Install:** `agentstack add mcp-annasmazhar-pyspark-mcp`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [AnnasMazhar](https://agentstack.voostack.com/s/annasmazhar)
- **Installs:** 0
- **Category:** [Databases](https://agentstack.voostack.com/c/databases)
- **Latest version:** 0.0.4
- **License:** MIT
- **Upstream author:** [AnnasMazhar](https://github.com/AnnasMazhar)
- **Source:** https://github.com/AnnasMazhar/pyspark_mcp

## Install

```sh
agentstack add mcp-annasmazhar-pyspark-mcp
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# PySpark MCP Server

SQL migration assistance, AWS Glue job generation, and Spark code optimization — as an MCP server.

[](https://github.com/AnnasMazhar/pyspark_mcp/actions/workflows/ci.yml)
[](https://www.python.org/downloads/)
[](https://opensource.org/licenses/MIT)

## What It Does

- **SQL Dialect Transpilation** — Convert between PostgreSQL, Oracle, Redshift, MySQL, Snowflake, and Spark SQL using [SQLGlot](https://github.com/tobymao/sqlglot)
- **PySpark DataFrame API Generation** — Generate DataFrame API code from SQL with optimization hints
- **AWS Glue Integration** — Job templates, DynamicFrame conversions, Data Catalog definitions, S3 optimization strategies
- **Batch Processing** — Process hundreds of SQL files concurrently
- **Code Review & Optimization** — Analyze existing PySpark code for performance improvements
- **Pattern Detection** — Find code duplication and suggest refactoring

## What It Doesn't Do

- Recursive CTEs → provides Spark SQL equivalent + guidance (PySpark has no native recursive CTE support)
- MERGE/PIVOT/CONNECT BY → transpiles to Spark SQL, provides DataFrame API guidance
- Perfect 1:1 DataFrame API transpilation for all SQL — complex queries get Spark SQL + optimization recommendations

## Quick Start

```bash
pip install -e .
pyspark-mcp  # starts the MCP server
```

## MCP Configuration

### Claude Desktop

Add to `~/Library/Application Support/Claude/claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "pyspark": {
      "command": "pyspark-mcp",
      "args": []
    }
  }
}
```

### Hermes Agent

Add to `~/.hermes/config.yaml`:

```yaml
mcp:
  servers:
    pyspark:
      command: pyspark-mcp
      enabled_tools: all
```

### Docker

```bash
docker compose up -d
```

## Tools

### SQL Conversion
- `convert_sql_to_pyspark` — Convert SQL to PySpark with dialect detection
- `analyze_sql_context` — Analyze SQL complexity and suggest approach

### AWS Glue
- `generate_aws_glue_job_template` — Generate complete Glue job scripts
- `convert_dataframe_to_dynamic_frame` — DataFrame ↔ DynamicFrame conversion
- `generate_data_catalog_table_definition` — Data Catalog table definitions
- `generate_incremental_processing_job` — Incremental/CDC job generation
- `analyze_s3_optimization_opportunities` — S3 layout and partitioning analysis

### Optimization
- `review_pyspark_code` — Code review with performance recommendations
- `optimize_pyspark_code` — Suggest optimizations for existing code
- `recommend_join_strategy` — Broadcast vs shuffle join recommendations
- `suggest_partitioning_strategy` — Partitioning recommendations

### Batch Processing
- `batch_process_files` — Process multiple SQL files concurrently
- `batch_process_directory` — Convert entire directories

## Development

```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

# Test
pytest tests/ -v --cov=pyspark_tools

# Format
black pyspark_tools tests
isort pyspark_tools tests

# Lint
flake8 pyspark_tools tests
```

## Architecture

```
pyspark_tools/
├── server.py              # FastMCP server + tool definitions
├── sql_converter.py       # SQLGlot-based transpilation + DataFrame API generation
├── aws_glue_integration.py # Glue job templates, DynamicFrame, Data Catalog
├── advanced_optimizer.py  # Performance analysis + optimization suggestions
├── batch_processor.py     # Concurrent file processing
├── code_reviewer.py       # PySpark code review patterns
├── duplicate_detector.py  # Code deduplication
├── data_source_analyzer.py # Data source analysis
└── file_utils.py          # File I/O utilities
```

## CI/CD

- ✅ 256 tests passing
- ✅ 71% code coverage
- ✅ Code quality checks (black, isort, flake8)
- ✅ Python 3.11 tested

## License

MIT — see [LICENSE](LICENSE).

---
`mcp-name: io.github.AnnasMazhar/pyspark-mcp`

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [AnnasMazhar](https://github.com/AnnasMazhar)
- **Source:** [AnnasMazhar/pyspark_mcp](https://github.com/AnnasMazhar/pyspark_mcp)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.0.4 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.0.4** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-annasmazhar-pyspark-mcp
- Seller: https://agentstack.voostack.com/s/annasmazhar
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
