Using MinerU
Quick Model Source Configuration
MinerU defaults to auto: probe Hugging Face first and choose ModelScope when it is unavailable. If users cannot access huggingface due to network restrictions, they can conveniently switch the model source to modelscope through environment variables:
export MINERU_MODEL_SOURCE=modelscope
Quick Usage via Command Line
MinerU has built-in command line tools that allow users to quickly use MinerU for document parsing through the command line:
mineru parse <input_path> -o <output_path>
Tip
<input_path>: One localPDF/OFD/EPUB/ staticHTML/MHTML(.mhtmlor.mht) / image /CSV/RTF/DOC/DOCX/PPT/PPTX/XLS/XLSX/ODT/ODS/ODPfile<output_path>: Optional output file; without it, Markdown is written to stdout- PDF parsing defaults to the first 10 pages; use
--pages allfor the full PDF. Non-PDF inputs, including MHTML, are parsed as whole documents and do not accept--pages.
For more information about output files, please refer to Output File Documentation.
Note
Runtime acceleration is selected separately for the two model components, based on installed dependencies and detected devices:
- Small models use the Torch backend only when
torch,torchvision,transformers,accelerate, andsafetensorsare all installed and a non-CPU device (CUDA/MPS/...) is detected; otherwise they run ONNX on CPU. Installing Torch alone is not enough. - The local VLM engine is chosen independently: macOS always uses llama.cpp; on an accelerator device Linux prefers vLLM, then an installed LMDeploy, and Windows uses LMDeploy; otherwise llama.cpp.
- XPU is excluded from automatic LMDeploy selection: Linux uses an installed XPU-compatible vLLM, otherwise llama.cpp; Windows uses llama.cpp.
- Windows users who need CUDA acceleration should first visit the PyTorch website and install accelerator-enabled
torchandtorchvisionmatching their CUDA version, then install themineru[full]extras.
After installation, confirm the runtimes actually in effect in your environment:
mineru-kit models show
The output reports Effective small backend and Effective VLM engine together with the config source of each value. models show displays effective backends, engines, and model readiness — missing model files are reported but do not fail the command; use mineru-kit models verify when a non-zero exit code should signal incomplete model files. Neither command runs inference. A full validation progresses in order: dependencies installed → selection as expected → models verify passes → a small sample document actually parses successfully. See Tiers and Runtimes for the full selection table.
If you need to adjust parsing options through custom parameters, you can also check the more detailed Command Line Tools Usage Instructions in the documentation.
Library, search, and continuation
mineru keeps file identity, parsing caches, and indexes in a local document library. Responses include locators; replace the example document ID below with an actual returned ID:
mineru parse document.pdf --json
mineru search "keyword" --json
mineru read "doc:ab12cd3/tier:standard/page:11" --json
When an output reaches its budget, follow next_request or the returned continuation command. read consumes existing results without automatically starting higher-quality parsing. Use mineru-kit parse for stateless batches and complete exports. Native Office, HTML/MHTML, CSV/TSV, EPUB, and OFD normalize to local Flash.
Advanced Usage via API, WebUI, and Services
-
Start the self-hosted V1 API:
mineru-kit api-server --host 0.0.0.0 --port 8000 --tier standardTip
Access
http://127.0.0.1:8000/docsfor the OpenAPI documentation. The supported service surface is/v1/*, including health, capability discovery, uploads, files, parse jobs, and usage.Native document parsing example: Python SDK; plain HTTP: V1 HTTP API
-
Start Gradio WebUI visual frontend:
mineru-kit webui --server-name 0.0.0.0 --server-port 7860Tip
- Access
http://127.0.0.1:7860in your browser to use the Gradio WebUI. - Without
--api-url, Gradio manages a loopbackmineru-kit api-server; with--api-url, it connects only to that existing V1 service. - Use
--api-server-preload-modelsto preload models for the managed local server. mineru-webuiremains available as a command-name alias with the same modern options.
- Access
-
Use
mineru-routerfor multi-service / multi-GPU orchestration:mineru-router --host 0.0.0.0 --port 8002 --local-gpus auto --worker-tier standardTip
mineru-routerandmineru-kit routerexpose the complete/v1/*API.- Repeat
--upstream-urlto aggregate multiple existing V1 api-server services, or use--local-gpusto launchmineru-kit api-serverworkers automatically. - Use
--preload-modelsfor router-managed workers; remote upstreams keep their own startup configuration. - Unknown model-engine arguments are not forwarded by Router.
- It is intended for advanced multi-service, multi-GPU, and unified-entry deployments.
-
Start an OpenAI-compatible VLM server:
mineru-kit vlm-server --engine auto --port 30000
Note
Model-engine parameters apply only to commands that explicitly declare them. mineru-router accepts documented Router/worker options and does not forward unknown arguments.
We have compiled some commonly used parameters and usage methods for vllm/lmdeploy, which can be found in the documentation Advanced Command Line Parameters.
Configuring LLM-aided post-processing with config.yaml
LLM-aided title leveling and cross-page table cell continuation read $MINERU_HOME/config.yaml and support OpenAI-compatible model services:
llm_aided:
api_key: ${MINERU_LLM_API_KEY:-}
base_url: https://dashscope.aliyuncs.com/compatible-mode/v1
model: qwen3.5-plus
enable_thinking: false
max_concurrency: 16
features:
title_leveling: false
cross_page_table_cell_merge: false
title_levelinggroups paragraph titles into levels 2 through 6 by document-title boundaries and runs only whenMiddleJson.is_full_documentistrue. Page-selected input is persisted asfalseand skips title leveling.cross_page_table_cell_mergeasks the LLM whether each pair of boundary-row cells continues after the existing rules identify a cross-page table.- Table cell merge does not require whole-document input. Both features are disabled by default and share one asynchronous client, one connection configuration, and the
max_concurrencyrequest limit, which defaults to 16. max_concurrencymust be an integer of at least 1 and can be overridden withMINERU_LLM_AIDED_MAX_CONCURRENCY.- Enabling either feature requires non-empty
api_key,base_url, andmodelvalues. enable_thinkingis optional. When omitted, the extension parameter is not sent to the model service.- The legacy
llm-aided-configsection inmineru.jsonis no longer read.
Extending MinerU Functionality with Configuration Files
MinerU works out of the box and reads current settings from $MINERU_HOME/config.yaml; set MINERU_CONFIG to use another file. Legacy mineru.json CLI settings are no longer supported.
Model storage and source settings use the model section:
model:
base_dir: ~/.mineru/models
source: auto
small_backend: auto
vlm:
engine: auto
See Model Source Documentation for model download and local-source details.
PDF page selection
Use --pages "1-5,8,r3-r1": page numbers start at 1, ranges include both endpoints,
and r1 means the last page. Use all for every page. Results are sorted and deduplicated;
partially out-of-bounds ranges select their valid intersection. Reversed or empty selections
fail with page_range_invalid. Without --pages, mineru parse starts with the first 10 pages;
mineru-kit parse, Python and Gradio select all pages. New requests use the current syntax. Historical positive result ranges using ASCII ~
remain readable without rebuilding Doclib caches; result responses and new cache entries use -.
Fullwidth ~ and negative page-number notation are not supported.
See page-range syntax and historical result compatibility.