Read-only local PDF MCP parsing. Extracted content enters your AI client and may reach its model provider.
Let your AI agent read a PDF without uploading it anywhere.
A Claude Code plugin that gives the agent local, read-only PDF tools (inspect, search, extract text) and a skill that checks a document for sensitive text before you share it. The server does not upload PDFs or modify them. Extracted text is returned to your AI client, which may send it to its model provider. Review that client's privacy settings and obtain consent before using sensitive documents.
- MCP server
pdfstudio, bundled inside the plugin (server/pdfstudio-mcp.mjs). It starts with--read-only, so the write tools (edit, redact, merge, split, rotate, page ops) are not exposed. - Read tools:
pdf_info,pdf_extract_text,pdf_search_text. - Skill
check-pdf-before-sharing: the agent searches the document for terms you name (a name, an account number, an address), pulls only the pages with hits, and reports each hit with its page and context.
Every read-tool result, including metadata, is marked as untrusted content; embedded delimiters are escaped. This helps the agent distinguish data from instructions, but is not a guarantee against prompt injection.
Plugin marketplaces are a Claude Code feature. From a Claude Code session:
/plugin marketplace add <path-or-git-url-to-this-folder>
/plugin install navigatorslab-privacy-tools@navigatorslab
Use the path to this folder, or the GitHub URL once it is published. The marketplace manifest is .claude-plugin/marketplace.json.
Requires Node 20+ on the machine that runs Claude Code.
node validate.mjsThe validator checks, without installing anything:
- the plugin manifest fields and name format
- that
.mcp.jsonpoints at the bundled server, is read-only, and has no absolute machine paths - that every skill's frontmatter name matches its folder and has a description
- that the server starts over stdio as configured and lists the read tools
- that a read-only server exposes no write tools
- real metadata, extraction and search calls against a bundled fixture
- rejection of existing outside-root PDFs through absolute paths, traversal and directory symlinks/junctions; huge page ranges, overlong search input and direct write-tool calls
- that fixture bytes remain unchanged
Exit 0 is valid, 1 lists the problems. A deliberately writable config fails the validator, which is the check that catches an accidental write-enabled setup.
The server is built from the PDF Studio MCP package and bundled with esbuild:
cd ../NavigatorsLab-PDF-Studio/packages/mcp && npm run build
cd - && npx esbuild ../NavigatorsLab-PDF-Studio/packages/mcp/dist/index.js --bundle --platform=node --format=esm --target=node20 --outfile=server/pdfstudio-mcp.mjsAlso ship pdfjs-dist/legacy/build/pdf.worker.mjs beside the server: PDF.js imports it at runtime even for local text extraction. Missing it makes extraction and search fail. The included synthetic fixture exercises this dependency.
Inputs are limited to 25 MiB; page ranges to 1,000 pages with page numbers at most 10,000; search queries to 1,000 characters; read responses to 50,000 characters. The chosen server root is the client's working directory (--root .), not a per-file consent sandbox. Use a dedicated document folder. Resource caps do not replace OS process memory/time limits or stop filesystem races caused by another local process.
- The server was started over stdio and returned its tool list.
- The validator passes on the shipped config and fails on a writable one.
- Not yet verified inside Claude Code. The install commands and
${CLAUDE_PLUGIN_ROOT}substitution follow Claude Code's plugin format, but this package has not been installed and exercised in a live session.
- Scanned PDFs without a text layer return no text. OCR is not part of this plugin.
- Search matches the text the PDF stores. Text drawn as images is invisible to it.
- Read-only means the agent cannot change the file. It does not stop the agent from quoting sensitive text in its reply, so review the output.
MIT
Use a supported Node LTS (22 or 24). Run npm ci --ignore-scripts followed by npm run verify. The security workflow runs the regression suite on Linux and Windows; it has not yet been executed remotely for this local change. New report/output files are created exclusively; choose fresh filenames or output folders for repeat runs. See SECURITY.md for the threat model and disclosure guidance.