This website requires JavaScript.
Explore
Help
Sign In
Mirror_Repos
/
ggml-org_llama.cpp
Watch
2
Star
0
Fork
0
You've already forked ggml-org_llama.cpp
Code
Issues
Pull Requests
Actions
Packages
Projects
Releases
Wiki
Activity
Files
d7cfe1ffe0f435d0048a6058d529daf76e072d9c
ggml-org_llama.cpp
/
tests
T
History
Johannes Gäßler
5fa07c2f93
CUDA: optimize FA for GQA + large batches (
#12014
)
2025-02-22 12:20:17 +01:00
..
.gitignore
…
CMakeLists.txt
…
get-model.cpp
…
get-model.h
…
run-json-schema-to-grammar.mjs
…
test-arg-parser.cpp
…
test-autorelease.cpp
…
test-backend-ops.cpp
…
test-barrier.cpp
…
test-c.c
…
test-chat-template.cpp
…
test-chat.cpp
…
test-double-float.cpp
…
test-gguf.cpp
…
test-grammar-integration.cpp
…
test-grammar-llguidance.cpp
…
test-grammar-parser.cpp
…
test-json-schema-to-grammar.cpp
…
test-llama-grammar.cpp
…
test-log.cpp
…
test-lora-conversion-inference.sh
…
test-model-load-cancel.cpp
…
test-opt.cpp
…
test-quantize-fns.cpp
…
test-quantize-perf.cpp
…
test-rope.cpp
…
test-sampling.cpp
…
test-tokenizer-0.cpp
…
test-tokenizer-0.py
…
test-tokenizer-0.sh
…
test-tokenizer-1-bpe.cpp
…
test-tokenizer-1-spm.cpp
…
test-tokenizer-random.py
…