NInfer vs llama.cpp vs vLLM: quality + speed comparison for Qwen3.8-27B NVFP4 on RTX 5090
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w821fg/ninfer_vs_llamacpp_vs_vllm_quality_speed/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA Reddit benchmark post comparing three inference engines (NInfer, llama.cpp, vLLM) on a Qwen3.8-27B NVFP4 model running on an RTX 5090, measuring both output quality and tokens-per-second.
Why it mattersHands-on comparison of inference backends on the newest consumer GPU with a fresh NVFP4 quantization format, useful for anyone sizing local LLM serving stacks.
Cited by
No citations on record.
