The first server deployment downloaded only 1101 bytes (a redirect/error
page) and the entrypoint promoted it to ggml-large-v3.bin, after which
every restart skipped the download and the server ran without a valid
model. Now both fresh downloads and existing files are validated
(minimum 50 MB + the ggml magic bytes); invalid files are logged,
deleted, and re-downloaded, and a failed download dumps the first 400
bytes of what was actually received for diagnosis before exiting.
Validated locally: a poisoned model file is detected, removed, and a
valid one re-downloaded; inference served correctly afterwards.
- server/whisper-server: compose stack built from the pinned whisper.cpp
v1.9.3 release, same as the Android JNI layer
- Vulkan GPU backend (AMD Radeon AI PRO R9700 / RADV) with transparent
CPU fallback and NO_GPU override; GGML models auto-download on first
start (MODEL env, default large-v3)
- API bound to the Tailscale interface only (100.103.83.12:8080) since
whisper-server has no authentication; render-group GID passthrough for
/dev/dri
- validated locally: image builds, entrypoint downloads tiny, POST
/inference returns verbose_json with language + segments; GPU-less
fallback confirmed