Part 1 covered what an engine is, and part 2 how to deploy one. The alternative is calling an API, where the provider picks the engine, the GPUs and the number formats. This post maps those options, what each provider says about its stack, and how to evaluate one you can’t inspect.

you control more, and operate more you operate less, and see less Run it yourself open engine on your GPUs YOU PICK engine, version, flags, GPUs, formats EXAMPLES vLLM, SGLang or TensorRT-LLM on your own cluster or cloud Managed open engine a platform runs it for you YOU PICK model, GPU type; often the engine EXAMPLES Baseten (picks among TRT-LLM, SGLang, vLLM), Anyscale (Ray + vLLM), Modal Proprietary stack their own engine on GPUs YOU PICK model, and a dedicated or shared deployment EXAMPLES Fireworks, Together Custom silicon their own chips, not GPUs YOU PICK model, from their list EXAMPLES Groq (LPU), Cerebras (wafer-scale)
Ways to serve a model. From left to right, you run less of the stack yourself, and you can check less of what runs it.

What providers say about their stacks Link to heading

Provider What they say Source
Baseten “we benchmark frameworks like TensorRT-LLM, SGLang, and vLLM to select the best-performing framework” guide, Jan 2026
Anyscale “Ray Serve for orchestration and scaling. vLLM for inference.” docs
DeepInfra “our inference stack is built on TensorRT-LLM and NVIDIA Dynamo” blog, Jun 2026
Fireworks “Fireworks proprietary LLM serving stack, which consists of CUDA kernels, optimized for both FP16 and FP8” blog, Jan 2024
Together the Together Inference Engine builds on “FlashAttention-3, faster GEMM & MHA kernels, innovations in quality-preserving quantization, and speculative decoding” blog, Jul 2024
Groq “Data flow is statically scheduled by the software during compilation” blog, Mar 2025
Cerebras “we are able to integrate 44GB of SRAM on a single chip” blog, Aug 2024

Fireworks’ current inference page says “We built it from the ground up, optimized every layer we control, from GPU memory layout to the runtime”. Whether any of it started from an open engine, neither Fireworks nor Together says.

The big labs say even less. Anthropic’s postmortem names its hardware, “AWS Trainium, NVIDIA GPUs, and Google TPUs”, but no engine. OpenAI names neither.

What you can’t see Link to heading

With a hosted API, these are someone else’s choices, and each one shows up in what you get:

  • Number format. An FP8 or FP4 deployment is cheaper to serve and can answer differently. Weight precision part 1 showed what that costs in accuracy.
  • Batching and scheduling. How busy the provider runs its GPUs sets your latency, especially the tail.
  • Engine version and settings. You don’t see them, and you don’t control when they change.

Reading provider benchmarks Link to heading

Proprietary stacks advertise against open engines, and the baselines age fast:

  • Together’s July 2024 post claims “up to 4.5x performance improvement over vLLM (version 0.5.1)”. vLLM is at 0.30 now.
  • Fireworks’ May 2025 post claims “3.5X throughput improvement compared to SGLang H200”, running on a B200: a newer GPU against an older one.

Neither claim is wrong. Both describe one setup on one date. Check the baseline version, the hardware on each side, and the date before using a number.

Evaluating a stack you can’t inspect Link to heading

Check How
Latency on your traffic Replay your own prompts at your own concurrency. Track first-token time and the p99 gap between tokens, as in Inference in production part 5.
Quality Run the same eval against the provider and against your own 16-bit baseline. A gap can mean a smaller number format.
The date on any comparison Ask which engine version and GPU the baseline used.
Capacity Shared endpoints have rate limits and noisy neighbors. Dedicated deployments cost more and behave more like your own.
Exit cost OpenAI-compatible APIs make switching cheap. Fine-tuned weights and provider-only features make it expensive.

Caveats Link to heading

  • Everything about proprietary stacks here is what the providers publish, with dates. None of it is verified independently.