Brandon Schneider
|
b54f597690
|
feat: ARM64 copy-if optimizer — branches to CSEL
Transforms branch patterns to ARM64 conditional selects:
Before: CMP + BEQ + compute + B + MOV = 5-47 cycles
After: CMP + compute + CSEL = 4-6 cycles
ARM64 CSEL instruction:
CSEL Xd, Xn, Xm, cond
- Single cycle on most ARM64 processors
- No branch prediction penalty
- No pipeline flush on mispredict
Pattern detection:
- CMP + BEQ/BNE/B.LT/etc
- True block: 1-3 compute instructions + B
- False block: single MOV
- Merge point
Same pattern as:
- SPIR-V OpSelect (GPU shaders)
- VCN delta+RLE (3.3x)
- QR spatial hash (2.18x)
- Lean CopyIfTactic (2.7x)
Works on ARM64 assembly from GCC/LLVM/Rust.
No compiler fork needed — post-processing pass.
Targets: Neon-64GB (18 vCPU ARM64 EPYC)
|
2026-05-30 15:59:24 -05:00 |
|