Part 1 covered what an engine is, and part 2 how to deploy one. The alternative is calling an API, where the provider picks the engine, the GPUs and the number formats. This post maps those options, what each provider says about its stack, and how to evaluate one you can’t inspect.
What providers say about their stacks Link to heading
| Provider | What they say | Source |
|---|---|---|
| Baseten | “we benchmark frameworks like TensorRT-LLM, SGLang, and vLLM to select the best-performing framework” | guide, Jan 2026 |
| Anyscale | “Ray Serve for orchestration and scaling. vLLM for inference.” | docs |
| DeepInfra | “our inference stack is built on TensorRT-LLM and NVIDIA Dynamo” | blog, Jun 2026 |
| Fireworks | “Fireworks proprietary LLM serving stack, which consists of CUDA kernels, optimized for both FP16 and FP8” | blog, Jan 2024 |
| Together | the Together Inference Engine builds on “FlashAttention-3, faster GEMM & MHA kernels, innovations in quality-preserving quantization, and speculative decoding” | blog, Jul 2024 |
| Groq | “Data flow is statically scheduled by the software during compilation” | blog, Mar 2025 |
| Cerebras | “we are able to integrate 44GB of SRAM on a single chip” | blog, Aug 2024 |
Fireworks’ current inference page says “We built it from the ground up, optimized every layer we control, from GPU memory layout to the runtime”. Whether any of it started from an open engine, neither Fireworks nor Together says.
The big labs say even less. Anthropic’s postmortem names its hardware, “AWS Trainium, NVIDIA GPUs, and Google TPUs”, but no engine. OpenAI names neither.
What you can’t see Link to heading
With a hosted API, these are someone else’s choices, and each one shows up in what you get:
- Number format. An FP8 or FP4 deployment is cheaper to serve and can answer differently. Weight precision part 1 showed what that costs in accuracy.
- Batching and scheduling. How busy the provider runs its GPUs sets your latency, especially the tail.
- Engine version and settings. You don’t see them, and you don’t control when they change.
Reading provider benchmarks Link to heading
Proprietary stacks advertise against open engines, and the baselines age fast:
- Together’s July 2024 post claims “up to 4.5x performance improvement over vLLM (version 0.5.1)”. vLLM is at 0.30 now.
- Fireworks’ May 2025 post claims “3.5X throughput improvement compared to SGLang H200”, running on a B200: a newer GPU against an older one.
Neither claim is wrong. Both describe one setup on one date. Check the baseline version, the hardware on each side, and the date before using a number.
Evaluating a stack you can’t inspect Link to heading
| Check | How |
|---|---|
| Latency on your traffic | Replay your own prompts at your own concurrency. Track first-token time and the p99 gap between tokens, as in Inference in production part 5. |
| Quality | Run the same eval against the provider and against your own 16-bit baseline. A gap can mean a smaller number format. |
| The date on any comparison | Ask which engine version and GPU the baseline used. |
| Capacity | Shared endpoints have rate limits and noisy neighbors. Dedicated deployments cost more and behave more like your own. |
| Exit cost | OpenAI-compatible APIs make switching cheap. Fine-tuned weights and provider-only features make it expensive. |
Caveats Link to heading
- Everything about proprietary stacks here is what the providers publish, with dates. None of it is verified independently.