e823bff873
Backport GGML kernels so we can enable flash attention for the gemma 4 model on Metal and CUDA.
Backport GGML kernels so we can enable flash attention for the gemma 4 model on Metal and CUDA.