Research-Stack/.hermes/plans/2026-05-29_220000-ray-deployment.md
Brandon Schneider e5fb0a5f4d chore: commit accumulated working tree changes
Lean: update Semantics modules, add new numerics/physics data files
Hardware: update FPGA bitstreams (tangnano9k_uart_loopback)
Infra: k3s-flake tests, netcup-vps configuration, VCN compute substrate
Docs: ARCHITECTURE, specs, citation updates
2026-05-30 00:10:02 -05:00

13 KiB

Plan: Ray Deployment on Research Stack k3s

Date: 2026-05-29 Status: Draft Estimated effort: 2-3 days


Goal

Deploy Ray on the Research Stack k3s cluster with:

  1. Distributed compute across 5 nodes (2 Ready, 3 flaky)
  2. H.264 transport via Ray Object Store
  3. GPU acceleration on qfox (RTX 4070)
  4. Thundering herd mitigations for node reconnection
  5. Graceful degradation when nodes go offline

Current State

Node Status Role Hardware
nixos Ready control-plane NixOS, 6.18.32 kernel
qfox-1 Ready worker CachyOS, RTX 4070
361395-1 NotReady worker Debian, edge VPS
racknerd NotReady worker Debian, VPS
steamdeck NotReady worker NixOS

Namespaces: ai-models (empty), services, media, mail, monitoring, research, edge

Existing workloads:

  • VCN pipeline: braid_vcn_encoder.py (Delta+RLE + RS ECC + ChaCha20 + H.264)
  • Fractal dimension: fractal_dimension.py (DBC algorithm, 29x vectorized)
  • GPU node: vcn_lupine_gpu_node.py (MKV decode + GPU compute)
  • AlphaProof: alphaproof_loop.py (Ollama → lake build → feedback)
  • Lean: 3572 jobs, 0 errors
  • Python: 68/68 tests pass

Architecture

Ray Topology

Neon-64GB (Ray Head + Worker)
├── GCS (Global Control Store) — PVC-backed
├── Dashboard (:8265) — via Traefik ingress
├── Object Store — 32GB allocated
├── Raylet — 18 ARM64 cores
└── VCN Encoder — Delta+RLE + RS ECC + ChaCha20

qfox-1 (Ray GPU Worker)
├── Raylet — RTX 4070
├── NVDEC — H.264 hardware decode (0.1ms)
├── CUDA — GPU compute (fractal dimension, Q16 LUT)
└── Object Store — 8GB allocated

nixos (Ray CPU Worker)
├── Raylet — CPU cores
├── Software decode — H.264 fallback
├── Lean builds — lake build via Ray tasks
└── Object Store — 4GB allocated

racknerd (Ray Edge Worker — standby)
├── Raylet — 2 vCPU EPYC
├── Edge relay — forward to external clients
└── Object Store — 2GB allocated

steamdeck (Ray Idle Worker — standby)
├── Raylet — idle until needed
└── Object Store — 2GB allocated

Network Flow

Braid data (Python)
    → Delta+RLE (vectorized copy-if, 3.3x)
    → RS ECC
    → ChaCha20
    → H.264 frame (YUV420)
    → MKV container
    → ray.put() → ObjectRef
    → Ray Object Store (shared memory)
    → GPU hardware decode (NVDEC) or CPU fallback
    → GPU compute or CPU compute
    → Result → ray.put() → ObjectRef
    → Downstream consumer

Port Map

Port Service Access
6379 Ray GCS Internal (Tailscale only)
8265 Ray Dashboard Via Traefik ingress
10001 Ray Client Tailscale only
8000 Ray Serve Via Traefik ingress

Implementation Steps

Phase 1: Ray Head Deployment (Day 1)

Step 1.1: Create Ray namespace and RBAC

  • File: 4-Infrastructure/k3s/ray/namespace.yaml
  • Create ai namespace (already exists, use it)
  • ServiceAccount, ClusterRole, ClusterRoleBinding for Ray

Step 1.2: Deploy Ray Head on Neon-64GB

  • File: 4-Infrastructure/k3s/ray/head-deployment.yaml
  • Image: rayproject/ray:2.43.0-py311-cu124
  • Command: ray start --head --dashboard-host=0.0.0.0 --object-store-memory=34359738368
  • PVC: 10Gi for GCS persistence
  • Resources: 4 CPU request, 16 CPU limit, 8Gi RAM request, 32Gi RAM limit
  • NodeSelector: kubernetes.io/hostname: neon-64gb
  • Liveness probe: ray health-check
  • Readiness probe: ray status

Step 1.3: Deploy Ray Dashboard Service

  • File: 4-Infrastructure/k3s/ray/dashboard-service.yaml
  • ClusterIP service on port 8265
  • Traefik IngressRoute at /ray/dashboard

Step 1.4: Verify Head is running

  • kubectl get pods -n ai
  • curl http://localhost:8265 (port-forward)
  • ray status inside head pod

Phase 2: Worker Deployment (Day 1)

Step 2.1: Deploy Ray GPU Worker on qfox-1

  • File: 4-Infrastructure/k3s/ray/gpu-worker-deployment.yaml
  • Image: rayproject/ray:2.43.0-py311-cu124
  • Command: ray start --address=ray-head.ai.svc.cluster.local:6379
  • Resources: 2 CPU, 8Gi RAM, 1 nvidia.com/gpu
  • NodeSelector: kubernetes.io/hostname: qfox-1
  • Toleration: nvidia.com/gpu:NoSchedule
  • Environment: NVIDIA_VISIBLE_DEVICES=all

Step 2.2: Deploy Ray CPU Worker on nixos

  • File: 4-Infrastructure/k3s/ray/cpu-worker-deployment.yaml
  • Image: rayproject/ray:2.43.0-py311-cu124
  • Command: ray start --address=ray-head.ai.svc.cluster.local:6379
  • Resources: 4 CPU, 8Gi RAM
  • No GPU resources

Step 2.3: Deploy Standby Workers (racknerd, steamdeck)

  • File: 4-Infrastructure/k3s/ray/standby-worker-deployment.yaml
  • Same as CPU worker but with lower resources
  • Toleration: node.kubernetes.io/unreachable:NoSchedule
  • Priority: low (evicted first when nodes are under pressure)

Step 2.4: Verify all workers connected

  • ray status inside head pod
  • Check: 4 workers connected (Neon, qfox, nixos, racknerd/steamdeck if online)

Phase 3: Thundering Herd Mitigations (Day 2)

Step 3.1: Staggered reconnection

  • File: 4-Infrastructure/k3s/ray/worker-entrypoint.sh
  • Script: compute jitter from Tailscale IP hash, sleep before connecting
  • delay = hash(tailscale_ip) % 30
  • Mount as ConfigMap, use as worker command

Step 3.2: Token bucket on Head

  • File: 4-Infrastructure/k3s/ray/head-entrypoint.sh
  • Configure Ray with --max-worker-startup-concurrency=1
  • One worker connects per second, burst of 3
  • Workers get 429 + Retry-After on rejection

Step 3.3: Object Store watermarks

  • File: 4-Infrastructure/k3s/ray/object-store-config.yaml
  • Configure Ray with memory thresholds:
    • 60%: throttle incoming syncs to 1/s
    • 80%: reject new connections
    • 95%: emergency GC
  • Spill cold objects to PVC at 60%

Step 3.4: Circuit breaker per worker

  • File: 4-Infrastructure/shim/ray_circuit_breaker.py
  • Python class wrapping Ray client calls
  • States: CLOSED (normal), OPEN (head overwhelmed), HALF-OPEN (testing)
  • Failure threshold: 5 failures in 30s → OPEN
  • Recovery: 60s wait → HALF-OPEN → probe → CLOSED

Step 3.5: Graduated task migration

  • File: 4-Infrastructure/shim/ray_task_migrator.py
  • When GPU node returns, migrate tasks in batches:
    • t=0s: 10% of tasks
    • t=30s: 20% more
    • t=60s: remaining 70%
  • Abort if any batch fails

Step 3.6: Witness deduplication

  • Ray ObjectRefs are already content-addressed
  • Verify: ray.put(data) returns same ObjectRef for same data
  • No custom code needed — document the behavior

Phase 4: VCN Pipeline Integration (Day 2)

Step 4.1: Wire VCN encoder to Ray Object Store

  • File: 4-Infrastructure/shim/vcn_ray_transport.py
  • encode_and_store(braid_data) → ObjectRef
  • decode_on_gpu(ObjectRef) → np.ndarray
  • decode_on_cpu(ObjectRef) → np.ndarray (fallback)

Step 4.2: Wire fractal dimension to Ray

  • File: 4-Infrastructure/shim/fractal_ray_task.py
  • @ray.remote(num_gpus=1) def compute_fd_gpu(data)
  • @ray.remote(num_cpus=4) def compute_fd_cpu(data)
  • Auto-select based on GPU availability

Step 4.3: Wire AlphaProof to Ray

  • File: 4-Infrastructure/shim/alphaproof_ray.py
  • @ray.remote def prove_theorem(problem)
  • Parallel proof search across workers
  • Pre-filter: skip trivial theorems (copy-if pattern)

Step 4.4: Test end-to-end

  • Encode 100 braid strands → Object Store → GPU decode → compute → verify
  • Measure: encode time, transfer time, decode time, compute time
  • Verify: 3.3x speedup on Delta+RLE, 100x on H.264 decode

Phase 5: k3s Migration (Day 3)

Step 5.1: Migrate control plane

  • Stop k3s-server on Neon-64GB
  • Install k3s-server on cupfox
  • Update kubeconfig to point to cupfox
  • Update Tailscale Funnel target

Step 5.2: Re-point workers

  • Update k3s-agent on all workers to point to cupfox
  • Verify: all nodes join new control plane

Step 5.3: Update DNS and ingress

  • Update researchstack.info A record → cupfox IP
  • Update Caddy upstream on racknerd
  • Verify: Traefik ingress works

Step 5.4: Verify Ray survives migration

  • Ray Head should reschedule on Neon (worker node now)
  • GCS persists on PVC (survives restart)
  • Workers reconnect to new Head

Phase 6: Monitoring and Observability (Day 3)

Step 6.1: Ray Dashboard ingress

  • File: 4-Infrastructure/k3s/ray/dashboard-ingress.yaml
  • Traefik IngressRoute at researchstack.info/ray/
  • SSO-gated via Authentik

Step 6.2: Prometheus metrics

  • Ray exports Prometheus metrics at :8080/metrics
  • Scrape config in monitoring namespace
  • Grafana dashboard for Ray cluster health

Step 6.3: Alerting

  • Alert on: worker disconnect, Object Store pressure, task failures
  • Route to: Telegram/Discord via Alertmanager

Files to Create

File Purpose
4-Infrastructure/k3s/ray/namespace.yaml Namespace + RBAC
4-Infrastructure/k3s/ray/head-deployment.yaml Ray Head on Neon
4-Infrastructure/k3s/ray/head-service.yaml ClusterIP + IngressRoute
4-Infrastructure/k3s/ray/gpu-worker-deployment.yaml GPU worker on qfox
4-Infrastructure/k3s/ray/cpu-worker-deployment.yaml CPU worker on nixos
4-Infrastructure/k3s/ray/standby-worker-deployment.yaml Standby on racknerd/steamdeck
4-Infrastructure/k3s/ray/worker-entrypoint.sh Staggered reconnection script
4-Infrastructure/k3s/ray/head-entrypoint.sh Token bucket config
4-Infrastructure/k3s/ray/object-store-config.yaml Watermark thresholds
4-Infrastructure/k3s/ray/dashboard-ingress.yaml Dashboard Traefik route
4-Infrastructure/shim/ray_circuit_breaker.py Circuit breaker per node
4-Infrastructure/shim/ray_task_migrator.py Graduated migration
4-Infrastructure/shim/vcn_ray_transport.py VCN → Ray Object Store
4-Infrastructure/shim/fractal_ray_task.py Fractal dimension on Ray
4-Infrastructure/shim/alphaproof_ray.py AlphaProof on Ray

Files to Modify

File Change
4-Infrastructure/k3s-flake/k3s-edge.nix Update Traefik for Ray ingress
6-Documentation/INFRASTRUCTURE.md Add Ray section
6-Documentation/RUNBOOK.md Add Ray troubleshooting
AGENTS.md Add Ray deployment rules

Validation

Phase 1 Validation

  • Ray Head pod Running on Neon
  • Dashboard accessible at localhost:8265 (port-forward)
  • ray status shows 1 node (Head)

Phase 2 Validation

  • GPU worker connected (qfox)
  • CPU worker connected (nixos)
  • Standby workers connected (racknerd, steamdeck if online)
  • ray status shows 4+ nodes

Phase 3 Validation

  • Simulate node disconnect: kill qfox pod
  • Verify: GPU tasks queued, not failed
  • Restart qfox pod
  • Verify: staggered reconnection (watch logs)
  • Verify: graduated migration (5→15→35→50 tasks)
  • Verify: Object Store watermark triggers at 60%

Phase 4 Validation

  • ray.put(braid_data) returns ObjectRef
  • ray.get(ObjectRef) returns original data
  • GPU decode works on qfox
  • CPU decode fallback works on nixos
  • 100 braid strands: end-to-end < 10s

Phase 5 Validation

  • Control plane migrated to cupfox
  • All workers join new control plane
  • Ray Head reschedules on Neon
  • GCS persists across restart

Phase 6 Validation

  • Dashboard accessible via Traefik
  • Prometheus scraping Ray metrics
  • Grafana dashboard shows cluster health

Risks and Mitigations

Risk Impact Mitigation
Neon ARM64 image compatibility Some services won't run Test images before deploying
Ray Head on 4GB cupfox OOM Keep Head on Neon (64GB), cupfox is control-plane only
Tailscale DERP latency Slow Ray RPC Use direct connections when possible
Object Store memory pressure OOM on workers Watermarks + PVC spill
GPU driver mismatch NVDEC fails Fall back to CPU software decode
etcd on cupfox 128GB disk Disk full Monitor disk usage, set alerts

Open Questions

  1. ARM64 Ray image: Does rayproject/ray:2.43.0-py311-cu124 have ARM64 builds? If not, need to build custom image.
  2. NVDEC on qfox: Is the NVIDIA driver new enough for NVDEC? Need to verify with nvidia-smi.
  3. PVC access from workers: Can Ray workers on different nodes mount the same PVC? Need ReadWriteMany or NFS.
  4. Ray version: 2.43.0 is latest stable. Check compatibility with k3s 1.35.4.
  5. Cost: Ray Head on Neon uses 8-32GB RAM. Is that acceptable alongside other workloads?

Success Criteria

  • Ray cluster operational with 4+ nodes
  • VCN pipeline runs on Ray (encode → transfer → decode → compute)
  • GPU acceleration working (NVDEC + CUDA)
  • Thundering herd mitigations tested (node disconnect/reconnect)
  • Dashboard accessible via Traefik
  • AlphaProof runs on Ray (parallel proof search)
  • All existing workloads still functional (k3s services, Lean, Python)