MarkItDown Skill logo

MarkItDown Skill

Visit

Agent skill for Microsoft MarkItDown 0.1.6: convert PDF, Office, HTML, CSV, EPUB, and ZIP files to Markdown for LLM and RAG use, with safe defaults.

Share:
View alternatives

MarkItDown: File to Markdown Conversion

The MarkItDown skill, maintained by K-Dense in Claude Scientific Writer (MIT), teaches an agent to use Microsoft's open-source MarkItDown library well. MarkItDown turns common documents into structure-preserving Markdown for search, text analysis, and LLM ingestion, not pixel-perfect reproduction. Version 2.0 of the skill targets MarkItDown 0.1.6 (released May 26, 2026) and corrects a common misconception: the built-in converters do not OCR scanned pages locally.

Key Features

  • Right method for the input: convert_local() for trusted files, convert_stream() with StreamInfo hints for bytes, and convert_response() after your own validated HTTP fetch. The permissive convert() and convert_uri() are reserved for trusted input.
  • Pinned install: uv pip install "markitdown[all]==0.1.6", or only the extras you need (pdf, docx, pptx, xlsx, xls, outlook, audio-transcription, youtube-transcription, and Azure extras).
  • Batch helpers: batch_convert.py converts a folder with a manifest, skips symlinks, and needs --allow-external-services before audio can go to a cloud transcriber. convert_literature.py converts paper PDFs with provenance front matter and can build an index.
  • OCR options: the official markitdown-ocr plugin (a vision-capable OpenAI-compatible client), Azure Document Intelligence, or Azure Content Understanding. None is installed by [all].
  • MCP server: markitdown-mcp exposes one tool, convert_to_markdown(uri); the skill recommends STDIO, and warns that HTTP/SSE mode has no authentication.

Use Cases

  • Preparing a folder of research PDFs and slide decks for a RAG index.
  • Extracting tables from XLSX or DOCX into Markdown an agent can read.
  • Giving Claude a local MCP tool for one-off document conversion.

Pricing and Access

MarkItDown and the skill are free and MIT-licensed. Local conversion can run offline. URL, YouTube, audio transcription (Google Web Speech via SpeechRecognition), LLM image descriptions, and Azure services send data outside your machine and may cost money.

Getting Started

  1. Install the plugin in Claude Code: /plugin marketplace add https://github.com/K-Dense-AI/claude-scientific-writer, then /plugin install claude-scientific-writer.
  2. uv venv --python 3.12 .venv, activate it, then install the pinned package.
  3. Run markitdown report.pdf -o report.md and compare headings and tables with the source.

Limitation: converted text is untrusted and can carry prompt injection. Scanned PDFs yield little text without an OCR plugin or Azure. Plugins run Python code in-process and are off by default.

Frequently Asked Questions

Why use result.markdown instead of text_content?

In 0.1.6 text_content is a soft-deprecated alias; new code should read result.markdown.

Does it keep page layout?

No. For bounding boxes or page coordinates the skill recommends a layout-aware parser instead.

Alternatives

  • PDF Skill: Anthropic's skill for merging, splitting, and creating PDFs, which MarkItDown does not do.
  • DOCX Skill: edits Word files rather than just reading them.
  • Firecrawl: a web-to-LLM crawling API for websites rather than local files.

Conclusion

The MarkItDown skill is a good default when an agent needs text out of many file types quickly and safely. Pin the version, keep processing local where possible, and verify output against the original. More document tools are in the skills hub.

Comments

No comments yet. Be the first to comment!