Home / Performance

llama.cpp · not a ranking

How fast will it run?

Pick a catalog model and a GPU chip. Results are llama.cpp CUDA decode and, when supported, prefill. Query combinations are not indexed. Methodology

Select a model and GPU. This page does not scrape benchmarks on request.