Skip to content

Using MinerU

Quick Model Source Configuration

MinerU defaults to auto: probe Hugging Face first and choose ModelScope when it is unavailable. If users cannot access huggingface due to network restrictions, they can conveniently switch the model source to modelscope through environment variables:

export MINERU_MODEL_SOURCE=modelscope
For more information about model source configuration and custom local model paths, please refer to the Model Source Documentation in the documentation.

Quick Usage via Command Line

MinerU has built-in command line tools that allow users to quickly use MinerU for document parsing through the command line:

mineru parse <input_path> -o <output_path>

Tip

  • <input_path>: One local PDF / OFD / EPUB / static HTML / MHTML (.mhtml or .mht) / image / CSV / RTF / DOC/DOCX / PPT/PPTX / XLS/XLSX / ODT/ODS/ODP file
  • <output_path>: Optional output file; without it, Markdown is written to stdout
  • PDF parsing defaults to the first 10 pages; use --pages all for the full PDF. Non-PDF inputs, including MHTML, are parsed as whole documents and do not accept --pages.

For more information about output files, please refer to Output File Documentation.

Note

Runtime acceleration is selected separately for the two model components, based on installed dependencies and detected devices:

  • Small models use the Torch backend only when torch, torchvision, transformers, accelerate, and safetensors are all installed and a non-CPU device (CUDA/MPS/...) is detected; otherwise they run ONNX on CPU. Installing Torch alone is not enough.
  • The local VLM engine is chosen independently: macOS always uses llama.cpp; on an accelerator device Linux prefers vLLM, then an installed LMDeploy, and Windows uses LMDeploy; otherwise llama.cpp.
  • XPU is excluded from automatic LMDeploy selection: Linux uses an installed XPU-compatible vLLM, otherwise llama.cpp; Windows uses llama.cpp.
  • Windows users who need CUDA acceleration should first visit the PyTorch website and install accelerator-enabled torch and torchvision matching their CUDA version, then install the mineru[full] extras.

After installation, confirm the runtimes actually in effect in your environment:

mineru-kit models show

The output reports Effective small backend and Effective VLM engine together with the config source of each value. models show displays effective backends, engines, and model readiness — missing model files are reported but do not fail the command; use mineru-kit models verify when a non-zero exit code should signal incomplete model files. Neither command runs inference. A full validation progresses in order: dependencies installed → selection as expected → models verify passes → a small sample document actually parses successfully. See Tiers and Runtimes for the full selection table.

If you need to adjust parsing options through custom parameters, you can also check the more detailed Command Line Tools Usage Instructions in the documentation.

Library, search, and continuation

mineru keeps file identity, parsing caches, and indexes in a local document library. Responses include locators; replace the example document ID below with an actual returned ID:

mineru parse document.pdf --json
mineru search "keyword" --json
mineru read "doc:ab12cd3/tier:standard/page:11" --json

When an output reaches its budget, follow next_request or the returned continuation command. read consumes existing results without automatically starting higher-quality parsing. Use mineru-kit parse for stateless batches and complete exports. Native Office, HTML/MHTML, CSV/TSV, EPUB, and OFD normalize to local Flash.

Advanced Usage via API, WebUI, and Services

  • Start the self-hosted V1 API:

    mineru-kit api-server --host 0.0.0.0 --port 8000 --tier standard
    

    Tip

    Access http://127.0.0.1:8000/docs for the OpenAPI documentation. The supported service surface is /v1/*, including health, capability discovery, uploads, files, parse jobs, and usage.

    Native document parsing example: Python SDK; plain HTTP: V1 HTTP API

  • Start Gradio WebUI visual frontend:

    mineru-kit webui --server-name 0.0.0.0 --server-port 7860
    

    Tip

    • Access http://127.0.0.1:7860 in your browser to use the Gradio WebUI.
    • Without --api-url, Gradio manages a loopback mineru-kit api-server; with --api-url, it connects only to that existing V1 service.
    • Use --api-server-preload-models to preload models for the managed local server.
    • mineru-webui remains available as a command-name alias with the same modern options.
  • Use mineru-router for multi-service / multi-GPU orchestration:

    mineru-router --host 0.0.0.0 --port 8002 --local-gpus auto --worker-tier standard
    

    Tip

    • mineru-router and mineru-kit router expose the complete /v1/* API.
    • Repeat --upstream-url to aggregate multiple existing V1 api-server services, or use --local-gpus to launch mineru-kit api-server workers automatically.
    • Use --preload-models for router-managed workers; remote upstreams keep their own startup configuration.
    • Unknown model-engine arguments are not forwarded by Router.
    • It is intended for advanced multi-service, multi-GPU, and unified-entry deployments.
  • Start an OpenAI-compatible VLM server:

    mineru-kit vlm-server --engine auto --port 30000
    

Note

Model-engine parameters apply only to commands that explicitly declare them. mineru-router accepts documented Router/worker options and does not forward unknown arguments. We have compiled some commonly used parameters and usage methods for vllm/lmdeploy, which can be found in the documentation Advanced Command Line Parameters.

Configuring LLM-aided post-processing with config.yaml

LLM-aided title leveling and cross-page table cell continuation read $MINERU_HOME/config.yaml and support OpenAI-compatible model services:

llm_aided:
  api_key: ${MINERU_LLM_API_KEY:-}
  base_url: https://dashscope.aliyuncs.com/compatible-mode/v1
  model: qwen3.5-plus
  enable_thinking: false
  max_concurrency: 16
  features:
    title_leveling: false
    cross_page_table_cell_merge: false
  • title_leveling groups paragraph titles into levels 2 through 6 by document-title boundaries and runs only when MiddleJson.is_full_document is true. Page-selected input is persisted as false and skips title leveling.
  • cross_page_table_cell_merge asks the LLM whether each pair of boundary-row cells continues after the existing rules identify a cross-page table.
  • Table cell merge does not require whole-document input. Both features are disabled by default and share one asynchronous client, one connection configuration, and the max_concurrency request limit, which defaults to 16.
  • max_concurrency must be an integer of at least 1 and can be overridden with MINERU_LLM_AIDED_MAX_CONCURRENCY.
  • Enabling either feature requires non-empty api_key, base_url, and model values.
  • enable_thinking is optional. When omitted, the extension parameter is not sent to the model service.
  • The legacy llm-aided-config section in mineru.json is no longer read.

Extending MinerU Functionality with Configuration Files

MinerU works out of the box and reads current settings from $MINERU_HOME/config.yaml; set MINERU_CONFIG to use another file. Legacy mineru.json CLI settings are no longer supported.

Model storage and source settings use the model section:

model:
  base_dir: ~/.mineru/models
  source: auto
  small_backend: auto
  vlm:
    engine: auto

See Model Source Documentation for model download and local-source details.

PDF page selection

Use --pages "1-5,8,r3-r1": page numbers start at 1, ranges include both endpoints, and r1 means the last page. Use all for every page. Results are sorted and deduplicated; partially out-of-bounds ranges select their valid intersection. Reversed or empty selections fail with page_range_invalid. Without --pages, mineru parse starts with the first 10 pages; mineru-kit parse, Python and Gradio select all pages. New requests use the current syntax. Historical positive result ranges using ASCII ~ remain readable without rebuilding Doclib caches; result responses and new cache entries use -. Fullwidth ~ and negative page-number notation are not supported. See page-range syntax and historical result compatibility.