Files
procleaning-mcp/README.md

60 lines
1.5 KiB
Markdown

# procleaning-mcp
Python MCP Server for PDF OCR batch processing.
## Features
- **Batch OCR** - Process multiple PDF files and save results as Markdown
- **Single File OCR** - Process a single PDF file directly
## Tool
### `batch_ocr_pdf`
OCR process PDF file(s) using the OCR API. Supports both single file and directory modes.
**Parameters:**
- `input_file`: Path to a single PDF file to process (mutually exclusive with `input_dir`)
- `input_dir`: Directory containing PDF files to process (mutually exclusive with `input_file`)
- `output_dir`: Directory to save Markdown results (auto-created if not provided)
- `api_key`: OCR API key (optional, reads from `OCR_API_KEY` env var)
- `base_url`: OCR API base URL (default: `https://agent.imqimacau.com`)
**Usage Modes:**
1. **Single File Mode**: Provide `input_file` only
```
input_file: /path/to/document.pdf
```
2. **Batch Mode**: Provide `input_dir` only
```
input_dir: /path/to/pdf_directory
```
**Note:** `input_file` and `input_dir` are mutually exclusive. At least one must be provided.
**Returns:** JSON with processing results (success/fail counts per file)
## Setup
```bash
cd procleaning-mcp
uv sync
uv run python -m procleaning_mcp.server
```
## Usage in Hermes
Add to `~/.hermes/config.yaml`:
```yaml
mcp_servers:
procleaning:
command: "uv"
args: ["run", "--project", "/home/tony/procleaning-mcp", "python", "-m", "procleaning_mcp.server"]
timeout: 300
```
Then use `mcp_procleaning_batch_ocr_pdf` in any Hermes conversation.