[ADD] odoo-at-payroll: wissensbasis skill + WIKU source integration model

Consolidate the binding knowledge-base rules (layer architecture,
licence decision D1, id freeze, D2 frontmatter schema, curation
conventions, batch workflow, validation gates, intake pitfalls)
from README/RUNBOOK/MEMORY into a dedicated agent skill so every
new thread reads them before planning Wissensbasis work.

- new licensed source corpus .wiku/ (WIKU Personal publications:
  Fachbroschueren, Arbeitsunterlagen, Casebooks, 'WIKU Personal
  aktuell'; gitignored like .lexis360/, never committed)
- integration model: one shared corpus with a single kb.json, ids
  wk-<prefix>-<nn> on the existing cluster map, the work field
  distinguishing the sources, Layer-2 filenames prefixed
  wiku_<slug>.md; pipeline to be a multi-source extension of
  build_lexis_kb.py (no fork) -- follow-up after batch 7
- AGENTS.md skill selection gains the wissensbasis entry
This commit is contained in:
2026-09-10 15:24:23 +02:00
parent 59bcda620f
commit fe8243931a
4 changed files with 248 additions and 0 deletions
+28
View File
@@ -21,6 +21,16 @@ dem Dev-Host bleibt das nächste Gate vor GP2 (Entgelt-Engine).
## Completed
- **Wissensbasis-Skill + WIKU-Quelle** (2026-09-10):
`.agents/skills/wissensbasis/SKILL.md` neu — verbindliche Regeln für
den Wissensbasis-Bau (Schichten/Lizenz D1, ID-Freeze, D2-Frontmatter,
Kurationskonventionen, Batch-Workflow, Validierungsgates, Intake-
Pitfalls) konsolidiert aus README/RUNBOOK/MEMORY; neu: WIKU-Quelle
`.wiku/` (31 lizenzierte PDFs: Fachbroschüren, Arbeitsunterlagen,
Casebooks, „WIKU Personal aktuell“) mit Integrationsmodell siehe
Entscheidung unten. `.wiku/` in `.gitignore` (D1); `AGENTS.md`
Skill-Auswahl ergänzt.
- **Lexis360-Wissensbasis Batch 6** (2026-09-10): 41 Briefings (Stände
2025-06 bis 2026-08) über zwei neue Kapitel (Arbeitnehmerschutz,
Verhaltenspflichten) + Schwerarbeit (Arbeitszeit-Kapitel) → 41
@@ -257,6 +267,20 @@ dem Dev-Host bleibt das nächste Gate vor GP2 (Entgelt-Engine).
beschaffen (anders als `.firecrawl/` = agenten-rekonstruierbare
Web-Fetches), sondern beim Nutzer nachfragen.
- **WIKU-Quelle `.wiku/`** (2026-09-10): D1 analog (lokal, unversioniert,
bei Fehlen NICHT beschaffen — Nutzer fragen). Integrationsmodell:
**gemeinsamer Korpus** `personalverrechnung/wissensbasis/` (eine
`kb.json`), `work`-Feld unterscheidet die Quellen; eigener ID-Raum
`wk-<prefix>-<nn>` auf der bestehenden Cluster-Map (`wk-pfa-01`
neben `lb-pfa-*`); Layer 1 `.wiku/md/` + eigener Katalog;
Layer-2-Dateinamen mit Präfix `wiku_<slug>.md`; Stand aus
„Stand YYYY-MM“/Heft-Monat, sonst Impressum; keine Lexis-Breadcrumbs —
Metadaten aus Dateiname + Titelei. Granularität: 1 Publikation =
1 Eintrag (Casebook-/Artikel-Ebene optional später); cross_refs lb↔wk
erwünscht. Pipeline = Multi-Source-**Erweiterung** von
`build_lexis_kb.py`, kein Fork — Folgeaufgabe **nach Batch 7**
(Lexis hat Priorität).
- **General-AT D1D4** (2026-09-09): Plan geführt als
`personalverrechnung/IMPLEMENTIERUNGSPLAN-Privat.md`; Referenz-KVs R1
= SI-2203 + SI-2748 (Feinjustierung je Pilotmandant in GP0);
@@ -297,6 +321,10 @@ dem Dev-Host bleibt das nächste Gate vor GP2 (Entgelt-Engine).
- `.lexis360/` — lizenzierte PDF-Exporte + Layer-1-Volltexte `md/` +
`_catalog.json` (lokal, unversioniert)
- `personalverrechnung/tools/build_lexis_kb.py` — KB-Pipeline
- `.agents/skills/wissensbasis/SKILL.md` — verbindliche Regeln für
Wissensbasis-Arbeit (inkl. `.wiku/`-Integrationsmodell)
- `.wiku/` — lizenzierte WIKU-Publikationen (lokal, unversioniert,
gitignored; Pipeline-Folgeaufgabe nach Batch 7)
- `personalverrechnung/IMPLEMENTIERUNGSPLAN-Privat.md` — General-AT
(freigegeben; GP0-Rechtsverifikationen erledigt, GP0-Rest offen;
+216
View File
@@ -0,0 +1,216 @@
---
name: wissensbasis
description: |
Rules for building and maintaining the curated Austrian personal-law
knowledge base (Wissensbasis) under personalverrechnung/wissensbasis/:
intake of licensed PDF sources (.lexis360/ Lexis Briefings Personalrecht,
.wiku/ WIKU Personal publications), Layer-1 extraction and cataloging via
personalverrechnung/tools/build_lexis_kb.py, Layer-2 curation (frontmatter
schema, clusters, frozen IDs, curation conventions, status marks),
registry generation (kb.json, INDEX.md), validation gates
(--registry/--check), the batch workflow, and licensing rules for the raw
sources. Use for any Wissensbasis work: new batches, curating entries,
pipeline/tool changes, or anything touching .lexis360/ or .wiku/.
disable-model-invocation: false
---
## Applicability
Use this skill for all work on the Wissensbasis: importing new batches,
curating Layer-2 entries, extending `build_lexis_kb.py`, regenerating
`kb.json`/`INDEX.md`, resolving corpus conflicts, and anything that reads
or writes the sources `.lexis360/` or `.wiku/`.
The Wissensbasis serves two purposes (see
`personalverrechnung/wissensbasis/README.md`):
1. **development reference** next to `RECHTSQUELLEN-*.md` for the payroll
modules (`l10n_at_hr_payroll*`), and
2. **future copilot corpus** — retrieval-ready: stable IDs, machine-readable
`kb.json`.
Skills do not replace the mandatory `AGENTS.md` workflow. When Wissensbasis
work feeds payroll implementation, combine with `payroll/SKILL.md`.
## Source corpora & licensing (decision D1, 2026-09-10)
| Source | Path | Content |
|---|---|---|
| Lexis 360 | `.lexis360/*.pdf` | licensed exports of *Lexis Briefings Personalrecht* |
| WIKU Personal | `.wiku/*.pdf` | licensed WIKU publications (Fachbroschüren, Arbeitsunterlagen, Casebooks, „WIKU Personal aktuell" issues) |
Binding rules:
- PDFs and extracted Volltexte are licensed content: **local + unversioned**
(both directories gitignored, pattern `.firecrawl/`). Never commit them.
- Only Layer-2 curation (own words, short quotes with source attribution)
is versioned.
- If a licensed source file is missing: **ask the user — never re-procure**,
never reconstruct from training knowledge. Unlike `.firecrawl/`, these
are not agent-reconstructable web fetches.
## Layer architecture
| Layer | Path | Versioned | Tool |
|---|---|---|---|
| PDF exports | `.lexis360/*.pdf`, `.wiku/*.pdf` | no | manual export/copy by the user |
| Layer 1 — full texts + catalog | `.lexis360/md/` + `_catalog.json`; WIKU: `.wiku/md/` (planned) | no | `--extract` |
| Layer 2 — curated entries | `personalverrechnung/wissensbasis/dokumente/<slug>.md` | **yes** | by hand |
| Registry + index | `personalverrechnung/wissensbasis/kb.json`, `INDEX.md` | **yes** (generated) | `--registry` |
The Layer-2 frontmatter is the **single source of truth**; `kb.json` and
`INDEX.md` are always regenerated from it, never hand-edited.
## IDs, clusters, frontmatter schema
- IDs: `lb-<prefix>-<nn>` (Lexis), `wk-<prefix>-<nn>` (WIKU). Assigned at
first `--extract`, then **frozen** (`load_previous_ids()` via
`_catalog.json`): never renumber, never reuse numbers of removed
documents. New documents append after the highest number in their
cluster.
- `topic` = descriptive ASCII cluster slug (`altersteilzeit`, `lehrlinge`,
…); the ID prefix lives only in `id`. New clusters extend **all four**
structures in `build_lexis_kb.py`: `TOPIC_MAP` (breadcrumb → prefix),
`KEYWORDS` (slug fallback), `CLUSTERS` (prefix → display name) and
`TOPIC_TO_PREFIX` (frontmatter validation).
- Decision D2: structural frontmatter keys in English (consistent with
`kv-catalog.json`/`chambers.json`), values in German UTF-8. Exceptions:
`stand` as ISO `YYYY-MM`, `topic`/`tags` as ASCII slugs (umlauts
**dropped**, not transliterated — `uberblick`; ß → `ss`).
- Layer-2 filename = Layer-1 slug (WIKU: with `wiku_` prefix, see below).
Binding frontmatter (full schema and cluster table:
`personalverrechnung/wissensbasis/README.md`):
```yaml
id: lb-atz-07 # frozen; WIKU: wk-<prefix>-<nn>
batch: 1 # procurement batch, set manually per import
title: "Altersteilzeit - Überblick"
work: "Lexis Briefings Personalrecht" # WIKU: exact publication name
chapter: "Beschäftigungsverhältnisse" # source chapter (breadcrumb; WIKU: derived)
topic: altersteilzeit # ASCII cluster slug
author: "Marek"
stand: 2026-01 # ISO month of the source's Stand
source:
pdf: ".lexis360/Lexis360_altersteilzeit_uberblick.pdf"
text: ".lexis360/md/altersteilzeit_uberblick.md"
legal_bases: ["AlVG", "AZG § 19e"] # only norms named in the source text
tags: [altersteilzeit, ams-foerderung] # ASCII slugs, specific before generic
cross_refs: ["lb-atz-09"] # related KB entries; dangling = error
```
## Curation conventions (binding)
1. **Sprache:** German, Fachsprache as in the original; metadata values
UTF-8; `topic`/`tags` ASCII.
2. **Eigene Worte** — curation is not a full-text copy (licence!). Short
verbatim quotes only, marked and with Stand.
3. **Werte immer mit Stand** — every value carries „(Stand YYYY-MM)".
Never add values from training knowledge — only from the source text,
or from newer KB entries (then cite the ID).
4. **Status marks as in RECHTSQUELLEN:** ✅ verified · ⚠ plausible, detail
verification open · ❓ deliberately open. Quote §§ only when the source
names them; otherwise ⚠ with a verification note (RIS).
5. **Document structure:** `# <Titel>` → *source line (work, author,
Stand, ID)* → `## Zusammenfassung``## Kernwerte & Fristen (Stand
YYYY-MM)` (table) → `## Rechtsgrundlagen``## Payroll-Relevanz
(Odoo)``## Verweise`.
6. **Payroll-Relevanz** names Odoo 19 anchor points (hr_payroll engine,
work entries, `hr.rule.parameter`, SV-BG handling, Meldewesen) as
implementation *hints*, not as a spec.
7. **Verweise:** KB IDs of related briefings (respect the source's
breadcrumb cross-references) + project files (`RECHTSQUELLEN-*.md`).
8. **Export artefacts:** ignore footers („Page n", „Erstellt von …");
never reconstruct truncated cross-references — note when a reference
spot is incomplete in the export.
9. **Reference / quality benchmark:** `dokumente/altersteilzeit_uberblick.md`.
## Batch workflow (new Lexis import)
```bash
# 1. copy new PDFs to .lexis360/ (keep export naming convention Lexis360_<slug>.pdf)
# 2. bump the BATCH constant in personalverrechnung/tools/build_lexis_kb.py
python3 personalverrechnung/tools/build_lexis_kb.py --extract # Layer 1 + catalog (IDs stay frozen)
# 3. curate new Layer-2 entries (conventions above)
python3 personalverrechnung/tools/build_lexis_kb.py --registry # kb.json + INDEX.md, validates frontmatter
python3 personalverrechnung/tools/build_lexis_kb.py --check # Layer-1<->Layer-2 completeness
```
Afterwards update `personalverrechnung/RUNBOOK.md` and `.agents/MEMORY.md`
(`INDEX.md` is regenerated, not hand-edited).
## Validation & values discipline
- `--registry` enforces: mandatory keys, ID pattern, ID-prefix↔topic
consistency, `stand` format, `batch` in `1..BATCH`, dangling
`cross_refs`.
- `--check` enforces: 1:1 catalog↔curation, Layer-1 text and PDF files
exist.
- Spot-check Kernwerte against the Layer-1 full text:
`sed -n '16,$p' .lexis360/md/<slug>.md`. Values never from training
knowledge.
- Corpus conflicts (source vs. source): document **both** values with IDs
and Stand in the entry (⚠/⚓) — never resolve silently.
`RECHTSQUELLEN-*.md` stays binding; RIS clarifies before implementation
(known conflicts: `personalverrechnung/RUNBOOK.md`, Wissensbasis
section, and `.agents/MEMORY.md`).
## Known intake pitfalls
- ß/URL-encoded export filenames (`%c3%9f` = ß) — normalize (ß → `ss`);
document it.
- Identical truncated export filenames for different briefings (Batch 7:
`auslandstatigkeit_sv_tatigkeit_in` ×3) — rename explicitly before
extraction.
- Duplicate exports with identical text — remove before extraction
(batches 2 and 4).
- Fossil IDs from Pass-1 keyword mis-grabs (e.g. `lb-mip-05` corrected to
`lb-swa-05`): fix before curation; the tool warns when a frozen ID has
a mismatching cluster prefix.
- Breadcrumb wrap variants: the parser must handle wrapping also after
the first „·" — harden for new variants when a batch surprises.
- `KEYWORDS` order matters (first match wins): **specific stems before
generic ones** (batch-4 lesson).
- Old Stände (e.g. leh 2024-03/2025-08, gsf/vst/lei 2025-06, son-0103
2025-06): curate only with explicit Stand marking.
## WIKU source `.wiku/` (integration model, 2026-09-10)
- Content: licensed WIKU Personal publications — Fachbroschüren,
Arbeitsunterlagen, Casebooks („gelöste Praxisfälle") and the periodical
„WIKU Personal aktuell" (issues „2026, Nr. N", combined issues like
„Nr. 4-5", „Nr. 8 - 9").
- Integration: **one shared corpus** — same
`personalverrechnung/wissensbasis/`, one `kb.json`; the `work` field
distinguishes the sources. WIKU entries use their own ID space
`wk-<prefix>-<nn>` on the existing cluster map (e.g. `wk-pfa-01`
alongside `lb-pfa-*`).
- Layer 1: `.wiku/md/<slug>.md` + `.wiku/md/_catalog.json` (separate
from Lexis). Layer-2 filenames get a `wiku_` prefix
(`wiku_lohnpfandung.md`) to avoid collisions in the shared `dokumente/`
directory.
- Stand determination: explicit „Stand YYYY-MM" / „YYYY-MM" in the
filename; periodicals: the issue's month; otherwise title page /
Impressum.
- WIKU has no Lexis breadcrumbs — metadata (Stand, chapter/topic) comes
from the filename and the title page.
- Granularity: **1 publication = 1 Layer-2 entry** (consistent with
briefing = entry). Per-case / per-article curation for Casebooks and
periodicals is a possible later extension.
- `cross_refs` between `lb-*` and `wk-*` on the same topic are
encouraged.
- Pipeline: WIKU must flow through a **multi-source extension of the
existing tool** (parameterize source directory and work), not a forked
second tool. Same validation gates apply. Status: not yet built —
follow-up **after Batch 7** (Lexis) is complete. Do not start WIKU
intake before that unless the user explicitly says so.
## Open points
- WIKU pipeline extension of `build_lexis_kb.py` (multi-source
`--extract`/`--registry`/`--check`).
- Copilot deployment needs a licence-compliant Layer-1 provisioning path
(Volltexte are not in the repo).
- Known corpus conflicts and verification backlogs: see
`personalverrechnung/RUNBOOK.md` (Wissensbasis section) and
`.agents/MEMORY.md` (open issues).