Home / Performance
llama.cpp · not a ranking
How fast will it run?
Pick a catalog model and a GPU chip. Results are llama.cpp CUDA decode and, when supported, prefill. Query combinations are not indexed. Methodology
Select a model and GPU. This page does not scrape benchmarks on request.