Planung und Skills für den Wissensbasis-RAG-Agenten

planung.md: Architektur (schlanker RAG-Service, SQLite-Index, Hybrid-Retrieval),
verbindliche Grounding-Regeln, Modell-Bake-off M3 (qwen3.8:27b, qwen3:32b,
gemma3:27b, mistral-small3.2:24b, qwen3:14b als Latenz-Untergrenze),
Meilensteine M1-M4 und Odoo-Integrationsoptionen.

.agents: neuer Skill pv-rag-agent (verbindliche Regeln für die Implementierung)
sowie bestehende Projekt-Skills (agent-memory, wissensbasis,
odoo19-development, opendataloader-pdf).
This commit is contained in:
2026-09-14 16:34:13 +02:00
parent 9d141eed9a
commit b5ca81abab
10 changed files with 2776 additions and 0 deletions
@@ -0,0 +1,176 @@
---
name: pv-rag-agent
description: |
Build and maintain the local RAG agent for Austrian payroll
(Personalverrechnung) that answers strictly from the curated
Wissensbasis (wissensbasis/, Layer 2, 601 entries). Covers the RAG
service pipeline (ingest, hybrid retrieval, generation, grounding,
eval), Ollama integration (server, models, VRAM budget), the binding
grounding and citation rules, the model bake-off protocol, and the
later Odoo Enterprise integration. Use for any work on the agent/
package, retrieval quality, prompts, the goldset/eval suite, Ollama
model choice, or the Odoo chat module. Combine with
wissensbasis/SKILL.md whenever Wissensbasis content changes.
disable-model-invocation: false
---
# PV RAG Agent (Wissensbasis-Copilot)
Implementation skill for the payroll knowledge agent planned in
`planung.md`. That file holds the full plan (architecture, milestones
M1M4, decision points); this skill records the rules a thread must
respect while implementing it.
## Scope & applicability
Use this skill for all work on:
- the RAG service (`agent/` package: ingest, retrieve, generate, api,
cli, eval) and its index (`data/index.db`);
- prompts, grounding checks, citation formatting, refusal behaviour;
- Ollama model configuration, embedding/reranker setup, server
connectivity;
- the eval goldset and any retrieval/prompt/model change;
- the later Odoo Enterprise chat module (Phase B).
When work touches the Wissensbasis itself (new batches, frontmatter,
curation), also apply `.agents/wissensbasis/SKILL.md`. When work
touches Odoo code, also apply `.agents/odoo19-development/SKILL.md`.
## System context (fixed facts)
- **Ollama server:** `http://100.183.83.12:11435` (custom port — do
not "correct" it to 11434). Verify reachability and installed models
with `curl http://100.183.83.12:11435/api/tags`. Note: agent sandboxes
may not reach this host — run such checks from the user's shell, not
the sandbox.
- **GPU:** AMD Radeon AI Pro R9700, 32 GB — keep the total resident
budget (answer model + embeddings + KV cache) under ~28 GB.
- **Models (provisional until bake-off, see protocol below):**
- answer model: `qwen3.8:27b` (Q4, ~18 GB, 256K context) — newest
Qwen generation (verified on the Ollama library 2026-09);
**thinking is on by default** — disable per request for RAG
latency (library documents per-request disabling plus
`reasoning_effort` / `preserve_thinking`; verify the exact Ollama
API option when implementing); vision exists but stays unused;
- fallback / known quantity: `qwen3:32b` (Q4_K_M, ~20 GB), thinking
mode **off**;
- further bake-off candidates: `gemma3:27b`, `mistral-small3.2:24b`
(24B, ~15 GB, 128K context; European vendor, expected strong
German — user hypothesis, verify in the bake-off; no thinking
mode), `qwen3:14b` (latency floor), `qwen3:30b-a3b` (MoE,
throughput);
- embeddings: `bge-m3` via `/api/embed` (multilingual, German);
- optional reranker: `bge-reranker-v2-m3` — verify the installed
Ollama version's rerank API **before** building on it; the design
must work without reranking (fallback: hybrid score only).
- **Corpus:** Layer 2 only — `wissensbasis/dokumente/*.md`, 601 entries,
frontmatter is the single source of truth; `kb.json` is its generated,
validated projection.
## Binding grounding constraints
These are the product's core promise — never weaken them:
1. **Answers only from the retrieved Layer-2 context.** No training
knowledge, no web search, no tools, no browsing hooks. The pipeline
has no outbound path besides Ollama — keep it that way.
2. **Citation duty:** every factual statement carries its KB ID
(e.g. `[lb-atz-07]`); every value carries its Stand
(`(Stand YYYY-MM)`), mirroring the Wissensbasis curation convention.
3. **Post-validation:** every ID cited in an answer must be in the
retrieved set. On violation: one regeneration with a stricter
instruction, then refuse or mark the answer as uncertain. Never ship
an answer that fails this check.
4. **Refusal duty:** if retrieval is empty or weak, say so ("Dazu
enthält die Wissensbasis keine Aussage") and optionally name related
clusters. Never fill gaps from prior knowledge.
5. **Corpus conflicts:** present **both** values with ⚠ and IDs (e.g.
ATZ replacement quota 28,5 % vs 27,5 %, lb-atz-07 vs lb-atz-09/12);
never resolve silently — same rule as curation convention 3.
6. **§ discipline:** cite norms only as the source names them
(`legal_bases`); no § completion from training knowledge.
7. **Layer 1 stays out of prompts** (`.lexis360/`, `.wiku/` are
licensed). Layer-2 text is curated own-words content and safe.
Layer-1 provisioning for deeper quotes is an open point (see
`wissensbasis/README.md`) — do not decide it ad hoc.
8. **Privacy:** the agent is a knowledge assistant. No employee or
payroll data flows into prompts — only the question and Layer-2
text.
## Architecture decisions (do not redesign without user approval)
- Lean custom pipeline, **no LangChain/LlamaIndex** (601 docs, full
control over grounding beats framework convenience).
- One SQLite file `data/index.db` (gitignored): FTS5 (BM25) + dense
vectors + metadata columns. No external vector DB.
- Chunking: H2 sections per entry; `## Kernwerte & Fristen` tables
become their own chunks (numeric questions); parent-child — retrieve
on section, provide section + metadata header as context.
- Hybrid retrieval: BM25 + dense (RRF fusion), optional reranker,
metadata filters (`topic`, `stand` recency), `cross_refs` expansion
of top hits; 812 context blocks, each with a metadata header
(ID · Titel · Stand · topic · Werk).
- FastAPI surface: `POST /ask`, `GET /health`, `POST /reindex`;
`agent/cli.py` for the dev loop; minimal static web UI for demos.
- German normalization for FTS5: umlaut folding at ingest time
(ä→ae or ä→a — pick once, stay consistent; ASCII slugs follow the
Wissensbasis convention: umlauts dropped, ß→ss).
## Ingestion rules
- Parse Layer-2 frontmatter directly from the `.md` files; treat
`kb.json` as a consistency gate (entry counts and ID sets must
match — mismatch aborts the ingest with a clear error).
- Incremental embeddings: cache vectors keyed by content hash; a
reindex only embeds new/changed chunks.
- After each new Wissensbasis batch (workflow in
`.agents/wissensbasis/SKILL.md`): run `POST /reindex`, then run the
eval suite.
## Validation gates
- **Goldset** `agent/eval/goldset.yaml`: 3050 questions with expected
IDs, including conflict cases (ATZ quotas) and 35 out-of-KB
questions that must be refused.
- **Metrics** (run before merging any retrieval/prompt/model change):
retrieval recall@8 (target > 0.9), citation precision (target 100 %),
refusal correctness, end-to-end latency.
- **Unit tests** (`tests/`): ingest schema validation, normalization,
post-validation behaviour (hallucinated ID → regeneration → refuse),
conflict rendering.
- Never validate with values from training knowledge — use the
Wissensbasis and Layer-1 spot checks instead.
## Model change protocol (bake-off, M3)
The answer model is only changed via a documented bake-off on the
goldset: `qwen3.8:27b` vs. `qwen3:32b` vs. `gemma3:27b` vs.
`mistral-small3.2:24b` (plus `qwen3:14b` as latency floor; temperature
~0.1, thinking off where the model has a thinking mode). Decision criteria: citation precision first, then refusal
correctness, then latency. Record the outcome like other project
decisions (D1/D2 style) in `planung.md` and `.agents/MEMORY.md`
(workflow: `.agents/SKILL.md` — agent-memory). A faster model may only
win if citation precision is equal.
## Odoo Enterprise integration (Phase B)
- **Option A (planned default):** thin custom module with an OWL chat
panel; service URL via `ir.config_parameter`; role-based access.
The RAG service remains the single source of truth for grounding and
citations. No retrieval/grounding logic in Odoo.
- **Option B (to verify first):** Odoo 19's native LLM modules with
Ollama as an OpenAI-compatible provider. **Never assume module
names, models, fields or endpoints** — verify against the actual
Odoo 19 source before planning Option B in detail
(`.agents/odoo19-development/SKILL.md`).
- Phase A runs independently of this decision; do not couple the
service API to Odoo specifics.
## Hygiene
- `data/index.db`, `data/`, logs and caches: gitignored.
- `.lexis360/`, `.wiku/`, `.firecrawl/`, `.ris/` stay unversioned
(licensed / local). Never commit them.
- Configuration via environment variables (`agent/config.py`): Ollama
URL, model names, port — no hardcoded hosts in business code.