Files
odoo_ocr/AGENTS.md
T
fegger 7e706793fa Initial scaffold: local invoice OCR pipeline with Ollama, classifier branches, structured extraction, review, and Odoo XML export
- Add AGENTS.md and project-specific Zed skills (odoo-ocr-pipeline, odoo-xml-import, local-vlm-client).
- Implement Pydantic schemas for documents, invoices, review results, and VLM responses.
- Add unified BaseVLMClient with Ollama implementation and llama.cpp stub.
- Build pipeline stages: loader, classifier, digital_pdf/scanned_print/handwritten/mixed_unknown branches, extractor, reviewer, xml_builder.
- Add CLI entry point  with sidecar JSON and confidence-gated XML output.
- Include prompts for classifier, OCR, extraction, and review models.
- Add tests with FakeVLMClient; pytest, ruff, and mypy all pass.
2026-08-21 14:04:42 +02:00

4.6 KiB

AGENTS.md — Odoo OCR Project

This file provides context and rules for AI agents working on odoo_ocr, a local-first invoice OCR pipeline that reads scanned, printed, handwritten, and native-PDF invoices and produces Odoo Enterprise-ready XML.

Project Goal

Build a Python application that:

  1. Classifies incoming invoice files by type (native PDF, scanned print, handwritten, mixed image).
  2. Routes each file to a specialized processing branch.
  3. Runs OCR using local vision-language models (Ollama or llama.cpp).
  4. Extracts a structured invoice representation.
  5. Reviews extracted data against the original image using a second vision model.
  6. Emits valid Odoo XML for import into Odoo Enterprise.

Tech Stack

  • Language: Python 3.11+
  • Dependency management: pyproject.toml (PEP 621); use uv or pip.
  • Core libraries:
    • pydantic v2 for all data schemas and settings.
    • pymupdf and pdf2image for PDF ingestion.
    • Pillow and opencv-python-headless for image preprocessing.
    • httpx for HTTP model clients.
    • lxml for XML generation and validation.
    • pytest and pytest-asyncio for tests.
  • Local inference:
    • Primary runtime: Ollama for fast iteration.
    • Optimized/runtime path: llama.cpp server (custom GGUF quants).
  • Models:
    • Document classifier: Qwen2.5-VL-3B or heuristics.
    • OCR / extraction: Qwen2.5-VL-7B or GLM-OCR (preferred for handwriting).
    • Review / validation: Qwen2.5-VL-7B or larger (72B if available).

Architecture

Input File
  → Classifier (heuristic + small VLM)
    → Branch: digital_pdf      → text/layout extraction
    → Branch: scanned_print    → preprocess → VLM OCR
    → Branch: handwritten      → preprocess → GLM-OCR / handwriting OCR
    → Branch: mixed_unknown    → preprocess → ensemble OCR
      → Structured Extraction VLM
      → Review VLM (image vs extracted JSON)
        → Confidence check
          → XML Builder
            → Odoo XML + sidecar JSON

Code Conventions

  1. Project layout: keep application code under src/odoo_ocr/.
  2. Schemas first: define Pydantic models before writing business logic.
  3. Type hints: use typing everywhere; run mypy in strict mode where practical.
  4. Error handling: never swallow exceptions; return structured Result objects or raise domain exceptions.
  5. Configuration: use pydantic-settings with config.yaml and env var overrides.
  6. Logging: use Python standard logging; log every model call latency and token usage.
  7. No hardcoded secrets: model endpoints, credentials, and paths come from settings.
  8. Tests: every module must have tests under tests/. Use fixtures from tests/fixtures/.

Model Client Rules

  1. Support both Ollama and llama.cpp with a unified interface (BaseVLMClient).
  2. Always emit JSON from VLMs when doing extraction/review. Use constrained prompts, not regex scraping.
  3. Retry on transient failures with exponential backoff.
  4. Cache model responses by content hash to avoid re-running expensive inference during development.
  5. Record per-call metrics (model name, tokens, latency, prompt hash).

Data Flow Rules

  1. Every invoice must produce:
    • A Pydantic ExtractedInvoice object.
    • A review result (ReviewResult) with confidence score and issue list.
    • An Odoo XML file (unless blocked by low confidence).
    • A sidecar JSON file with metadata, timings, and confidence.
  2. If review confidence is below the configured threshold, mark the invoice for human review and do not generate final XML (or generate a draft with a warning flag).
  3. Never send invoice data outside the local model endpoints.

Odoo XML Target

Generate Odoo data-import XML compatible with Odoo Enterprise vendor bills:

  • res.partner (vendor)
  • account.move (vendor bill header)
  • account.move.line (invoice lines)
  • account.tax references by percentage/name

See skill odoo-xml-import for detailed field mapping and a sample XML template.

When to Ask the User

Ask for clarification when:

  • The requested change would alter the model stack.
  • The change affects the Odoo target schema or import method.
  • You are unsure whether a file should be committed or a dependency added.
  • A requested feature conflicts with the local-only / privacy constraint.

Skills Reference

Agents should load the relevant project skills from .agents/skills/:

  • odoo-ocr-pipeline — when implementing classifier, branches, OCR, extraction, or review.
  • odoo-xml-import — when generating or validating Odoo XML.
  • local-vlm-client — when writing model clients or prompts.