Commit graph

4 commits

Author SHA1 Message Date
Brandon Schneider
a031a73dfe fix(infra): standardize k3s cluster — Traefik, ingress, architecture docs, Ray GPU workers
Architecture alignment:
- Rewrite k3s-server.nix from aspirational role=server to actual role=agent
  (cupfox is the real control-plane at 100.110.163.82:6443)
- Update flake.nix: correct serverAddr for all nodes, annotate dead nodes,
  fix hostname from nixos-laptop to nixos
- Update join-agent.sh default SERVER to cupfox
- Document actual vs intended architecture in comments

Ingress:
- Add rs-apps-books Ingress for audiobookshelf (/apps/books → media ns)
- Add strip-apps-books middleware for prefix stripping
- Create ray-ingress.yaml for Ray dashboard at /server/ray

Ray:
- Fix raycluster.yaml: enable dashboard, mount /dev/dri on gpu-workers
- Point gpu-workers to qfox-1 (was neon-64gb)
- Remove nvidia.com/gpu resource dependency (use /dev/dri via Mesa)

Security:
- Move OPENID_CLIENT_SECRET from plaintext to K8s Secret ref
- Update manifest to use valueFrom.secretKeyRef

Cleanup:
- Remove 53 committed test PNG screenshots from git tracking
- Remove auth-state.json from git tracking
- Add *.png, *.json to tests/.gitignore

Build: no build needed
2026-05-31 23:13:04 -05:00
Brandon Schneider
10670e2d10 feat(infra): Ray VCN transport + cluster restoration
- ray_vcn_transport.py: @ray.remote wrappers for braid VCN encode/decode
  - Distributed encode on CPU workers, compute on GPU workers
  - RayVCNTransport actor with frame counter + ObjectRef storage
  - FAMM-gated encode task, batch encode/decode helpers
  - 20 strands in 576ms (28.8ms/strand), 20/20 CRC ok

- raycluster.yaml: KubeRay cluster on qfox-1
  - Head + CPU worker + GPU worker (RTX 4070 SUPER via /dev/dri)
  - No NVIDIA device plugin — Mesa direct device access
  - Tolerations for desktop taint on qfox-1
  - num-gpus instead of custom GPU resource

- fix-nftables-k3s.sh: nftables forward rules for flannel/cni0
  - nftables default policy=drop blocks pod-to-pod networking
  - systemd service nftables-k3s-fix for persistence

- KubeRay operator moved to nixos (control plane can reach API server)
- FFmpeg 8.0 + reedsolo installed in Ray head pod via conda
2026-05-30 19:42:39 -05:00
Brandon Schneider
c6206b1ba8 docs(kube): update infra docs, RayCluster manifest with nightly GPU images
- infrastructure-status.md: rewrite with current cluster topology (cupfox control-plane,
  neon-64gb/racknerd/steamdeck workers), RayCluster status, Garage storage,
  Caddy edge, open issues
- k3s-cluster-setup.md: fix steamdeck hardware specs (8 vCPU, 14.5 GB RAM)
- raycluster.yaml: upgrade to rayproject/ray:nightly-py313-gpu (multi-arch amd64+arm64),
  add gpu-workers group targeting neon-64gb, add arm64-workers for neon-64gb CPU
- README.md: update build job count (3460 → 3313, verified)

Build: lake build Compiler 3313 jobs, 0 errors
2026-05-30 16:23:13 -05:00
Brandon Schneider
25f0ec2b53 feat: QR spatial hash integration — 2.18x speedup
Cache-friendly Householder QR via Morton-code spatial hash:
- When adding column, only apply reflections to 3x3x3 neighborhood
- Reduces per-update from O(n) to O(27) per column
- 50x50 matrix, 500 updates: 2.18x faster than naive

Naive: 0.124ms/update
Spatial: 0.057ms/update
Speedup: 2.18x

Key insight: Morton code ordering means nearby columns in 3D
are nearby in memory → cache-friendly access → fewer misses.

This completes all 4 next steps:
1.  O_AMMR_QRNode wired into BraidDiatFrame (already done)
2.  O_AMMR_valid strengthened with residual bounds (NS_MD.lean)
3.  Hash benchmark: Morton wins (86.5% cache hit rate)
4.  QR spatial hash: 2.18x speedup
2026-05-30 15:30:06 -05:00