- Ollama: qwen3.5:4b-fast with NUM_PARALLEL=8, num_ctx=16384 (~6GB VRAM) - vLLM: numind/NuExtract3 at 192.168.42.254:2701, gpu-memory-util=0.45 - All 3 pipelines (beta/paper/live) updated with correct endpoints - Ollama proxy at 10.1.1.12:2701 (K8s ollama-metrics) - vLLM proxy at 192.168.42.254:2701 (NixOS nginx)