- ray_vcn_transport.py: @ray.remote wrappers for braid VCN encode/decode
- Distributed encode on CPU workers, compute on GPU workers
- RayVCNTransport actor with frame counter + ObjectRef storage
- FAMM-gated encode task, batch encode/decode helpers
- 20 strands in 576ms (28.8ms/strand), 20/20 CRC ok
- raycluster.yaml: KubeRay cluster on qfox-1
- Head + CPU worker + GPU worker (RTX 4070 SUPER via /dev/dri)
- No NVIDIA device plugin — Mesa direct device access
- Tolerations for desktop taint on qfox-1
- num-gpus instead of custom GPU resource
- fix-nftables-k3s.sh: nftables forward rules for flannel/cni0
- nftables default policy=drop blocks pod-to-pod networking
- systemd service nftables-k3s-fix for persistence
- KubeRay operator moved to nixos (control plane can reach API server)
- FFmpeg 8.0 + reedsolo installed in Ray head pod via conda
Cache-friendly Householder QR via Morton-code spatial hash:
- When adding column, only apply reflections to 3x3x3 neighborhood
- Reduces per-update from O(n) to O(27) per column
- 50x50 matrix, 500 updates: 2.18x faster than naive
Naive: 0.124ms/update
Spatial: 0.057ms/update
Speedup: 2.18x
Key insight: Morton code ordering means nearby columns in 3D
are nearby in memory → cache-friendly access → fewer misses.
This completes all 4 next steps:
1. ✅ O_AMMR_QRNode wired into BraidDiatFrame (already done)
2. ✅ O_AMMR_valid strengthened with residual bounds (NS_MD.lean)
3. ✅ Hash benchmark: Morton wins (86.5% cache hit rate)
4. ✅ QR spatial hash: 2.18x speedup