570c8585a04e7f435315fd8c9d61db92f8519dea
- Replaces duplicated _load_prompt() helpers with shared load_prompt(). - Classifier raises cleanly when no renderable pages exist. - scanned_print, handwritten, and mixed_unknown branches now force response_format=OcrResponse and parse via client.parse_json(). - Adds branch tests for multi-page OCR and corrected invoice handling.
odoo_ocr
Local-first invoice OCR pipeline that reads scanned, printed, handwritten, and native-PDF invoices and produces Odoo Enterprise-ready XML.
Quick start
-
Install dependencies:
uv sync --all-extras # or pip install -e ".[dev]" -
Configure
config.yamlor set environment variables:export OLLAMA_BASE_URL="http://100.103.83.12:11435" -
Run the pipeline:
odoo-ocr process /path/to/invoices --output ./out/
Architecture
Input File
→ Classifier (heuristic + small VLM)
→ Branch: digital_pdf → text/layout extraction
→ Branch: scanned_print → preprocess → VLM OCR
→ Branch: handwritten → preprocess → GLM-OCR
→ Branch: mixed_unknown → preprocess → ensemble OCR
→ Structured Extraction VLM
→ Review VLM (image vs extracted JSON)
→ Confidence check
→ XML Builder
→ Odoo XML + sidecar JSON
Project-specific agent skills
Agent skills are in .agents/skills/:
odoo-ocr-pipeline— classifier, branches, extraction, review.odoo-xml-import— generating and validating Odoo XML.local-vlm-client— Ollama/llama.cpp VLM clients and prompts.
Description
Languages
Python
99.5%
Shell
0.5%