Thirteen Llama hosting providers were compared on price, context, compliance and access, collected September 2, 2026.
DeepInfra posts the lowest Llama 3.3 70B output price, $0.32 per 1M tokens, 73% below SambaNova Cloud's $1.20. Groq, Cerebras and Lambda each cut self-serve Llama access this year. DeepInfra or Fireworks AI settle general use; read on only for a specific constraint.
Llama Hosting Providers Compared
Provider | Llama versions confirmed | Flagship Llama price | Best context delivered | Compliance confirmed | Free tier | Verdict |
|---|---|---|---|---|---|---|
DeepInfra | 3.1 8B/70B, 3.3 70B, 4 Scout, 4 Maverick | $0.10 / $0.32 per 1M (3.3 70B) | 1.02M tokens (4 Maverick) | SOC 2, HIPAA, GDPR in progress | Not confirmed | Top pick, cost |
Fireworks AI | 3.1 8B/70B/405B, 3.2 1B/3B/11B/90B, 3.3 70B, 4 Maverick | $0.90 / $0.90 per 1M (size-band rate) | 1.05M tokens (4 Maverick) | SOC 2, HIPAA, GDPR mapped | $1 credit | Top pick, compliance and lineup |
Together AI | 3.3 70B, 4 Maverick | $1.04 / $1.04 (3.3 70B); $0.27 / $0.85 (4 Maverick) | 1.05M tokens (4 Maverick) | SOC 2, HIPAA | None, $5 minimum purchase | Close second |
AWS Bedrock | 3.2 1B/3B/11B, 3.3 70B, 4 Scout, 4 Maverick | $0.72 / $0.72 (3.3 70B) | 3.5M tokens (4 Scout) | Not re-verified this session | Not confirmed | Situational, longest context |
Novita AI | 3.1 8B, 3.3 70B, 4 Scout, 4 Maverick | $0.135 / $0.40 (3.3 70B, capped at 12K context) | 1M tokens (4 Maverick) | SOC 2 badge only | Voucher, amount undisclosed | Situational, price with a context catch |
SambaNova Cloud | 3.3 70B only | $0.60 / $1.20 | 128K tokens | Not confirmed, gated trust center | Not confirmed | Insufficient data |
Hyperbolic | 3.1 405B (deprecating), 3.3 70B | $0.40 blended (3.3 70B) | 131K tokens | GDPR only, SOC 2 pursuing, HIPAA denied | Not confirmed | Situational, EU residency |
Groq | 3.1 8B, 3.3 70B | Enterprise, contact sales | 131K tokens | HIPAA (BAA), SOC 2 not confirmed | Not confirmed | Situational, fastest published tok/s |
Baseten | 3.1 8B/70B/405B, 3.2 11B/90B, 3.3 70B, 4 Scout, 4 Maverick | No per-token price, GPU-hour only ($4.00 to $6.50/hr) | Not published | SOC 2, HIPAA | Undisclosed credits | Insufficient data, dedicated only |
Replicate | Llama 3 8B/70B, 4 Scout, 4 Maverick | $0.25 / $0.95 (4 Maverick) | Not published for 4 Maverick | Not confirmed | Not confirmed | Insufficient data |
RunPod | Any version, self-managed | $1.59/hr (A100 80GB), $3.29/hr (H100) | User-configured | SOC 2, SOC 3, HIPAA and GDPR programs | Not confirmed | Top pick, self-hosting |
Lambda | Any version, self-managed | $2.79/hr (A100 80GB), $3.99/hr (H100) | User-configured | SOC 2, ISO 27001/27017/27701/22301 | Not confirmed | Close second, self-hosting |
Cerebras | None | Not applicable | Not applicable | Not applicable | $5 credit (non-Llama models) | Not recommended |
Rows are grouped by pricing model: per-token providers first, ranked by lowest confirmed output price; GPU-rental providers next, ranked by A100 rate; Cerebras last, since it hosts no Llama model at all. A dash means the figure was not published on a page this article's research fetched directly, not that the feature is absent.
DeepInfra prices Llama 3.3 70B output at $0.32 per 1M tokens against SambaNova Cloud's $1.20, a 73% gap on the same model.
Fireworks AI is the only provider confirmed on SOC 2 Type II, HIPAA and GDPR-mapped controls together, alongside the widest confirmed Llama lineup: eight model sizes across three generations.
Basis and method
No figures in this comparison were produced by testing. Every price, context limit, compliance claim and free-tier detail was collected directly from each provider's own pricing page, API documentation or trust center during a single session on September 2, 2026, and every source is recorded with its URL in the raw data file and on the LLM Waves methodology page.
Two derived datasets appear below: a VRAM sizing table computed from Meta's published parameter counts, and a self-hosting breakeven calculation built from RunPod's rental rate and a GPU vendor's own published throughput figure, both formulas published inline.
Where a trust center failed to load or returned a gated page (SambaNova, Hyperbolic's signup flow), that failure is noted at the point it occurred rather than repeated as a blanket disclaimer.
No figure below was independently verified beyond confirming that the cited page loaded and displayed the quoted text.
Decision axes
Providers are ranked on four axes: per-token price for a Llama model both providers host, the context window actually delivered against Meta's stated native window, the number of compliance certifications confirmed from a primary source, and whether self-serve Llama access is stable or has been withdrawn.
The full, continuously refreshed version of this comparison lives on the LLM Waves inference data pages; the tables below are a snapshot of it.

DeepInfra's Llama 3.3 70B pricing is the floor of the group at $0.10 input and $0.32 output per 1M tokens, collected from DeepInfra's own model page on September 2, 2026.
SambaNova Cloud prices the same model at $0.60 input and $1.20 output, a 73% gap on output that holds regardless of which other provider in the table gets compared against it.
Four providers, Hyperbolic, Fireworks AI, AWS Bedrock and Together AI, price input and output identically, which flattens any advantage from prompt-heavy workloads.
Which Llama version does each provider actually support?
Meta ships five active Llama generations: Llama 3.1 at 8B, 70B and 405B parameters, Llama 3.2 at 1B, 3B, 11B and 90B, Llama 3.3 at 70B, and Llama 4 Scout and Maverick, both mixture-of-experts models.
A provider's homepage claim of "Llama support" rarely states which of these five it actually ships, and the matrix below states it directly, provider by provider.
Provider | 3.1 | 3.2 | 3.3 70B | 4 Scout | 4 Maverick |
|---|---|---|---|---|---|
Fireworks AI | 8B, 70B, 405B | 1B, 3B, 11B, 90B | Yes | Not confirmed | Yes |
Baseten | 8B, 70B, 405B | 11B, 90B | Yes | Yes | Yes |
DeepInfra | 8B, 70B | Not confirmed | Yes | Yes | Yes |
AWS Bedrock | Not confirmed | 1B, 3B, 11B | Yes | Yes | Yes |
Novita AI | 8B | Not confirmed | Yes | Yes | Yes |
Together AI | Not confirmed | Not confirmed | Yes | Not confirmed | Yes |
Groq | 8B (enterprise) | Not confirmed | Yes (enterprise) | Not confirmed | Not confirmed |
Hyperbolic | 405B (deprecating) | Not confirmed | Yes | Not confirmed | Not confirmed |
SambaNova Cloud | Not confirmed | Not confirmed | Yes, only model | Not confirmed | Not confirmed |
Replicate | Llama 3, not 3.1 | Not confirmed | Not confirmed | Yes | Yes |
Cerebras | Deprecated | Not applicable | Deprecated | Deprecated | Deprecated |
RunPod, Lambda | Self-managed, any version, GPU-rental model rather than a fixed catalog |
"Not confirmed" means this research did not find that model size on the provider's own pricing or model page, not that the provider has ruled it out. Baseten and Fireworks AI cover the widest range: eight variants each across three generations. Groq restricts both of its listed models, llama-3.1-8b-instant and llama-3.3-70b-versatile, to an enterprise tier priced by contact rather than a published rate, confirmed on Groq's own model documentation.
Context window: what Meta announces and what you actually get
Meta's own Llama 4 announcement states a native context window of 10 million tokens for Llama 4 Scout and 1 million tokens for Llama 4 Maverick.
Neither figure survives contact with a hosting provider's actual limit.

The widest Llama 4 Scout window found on a provider's own page in this research is 3.5 million tokens, offered by AWS Bedrock and confirmed on the AWS News Blog's Bedrock announcement, a 65% reduction from Meta's stated 10 million.
DeepInfra caps the same model at 320,000 tokens, a 97% reduction. Llama 4 Maverick fares better: Together AI, DeepInfra and Novita AI all deliver 1,048,576 tokens, matching Meta's 1 million-token claim almost exactly.
Llama 3.1 and Llama 3.3 both hold their advertised 128,000-token window everywhere this article found a stated limit, with one exception: Novita AI caps Llama 3.3 70B at 12,000 tokens on its own pricing table, a limit worth checking before assuming any 70B endpoint gives the full 128K window.
The quantization and serving choices behind gaps like this one are explained in why the same model performs differently across providers.
Provider access is not fixed: three vendors pulled back within a year
A hosting comparison built once and left unrefreshed goes stale fastest on this axis, and the research for this article turned up three examples inside a twelve-month window.
Cerebras's own deprecation log shows every Llama model it once served removed on a rolling schedule: Llama 4 Maverick on October 15, 2025, Llama 4 Scout on November 3, 2025, Llama 3.3 70B on February 16, 2026, and the last remaining model, Llama 3.1 8B, on May 27, 2026. Cerebras's current public catalog lists only GPT-OSS and Gemma.
Groq moved both of its Llama models behind an enterprise sales gate, confirmed on its own model documentation, replacing public per-token pricing with a contact-sales listing.
Lambda's own inference page states plainly that "as the Inference API winds down, you can continue deploying and scaling models seamlessly on NVIDIA GPU instances," which converts Lambda from a managed Llama API into a GPU-rental platform with no managed Llama product at all.
None of these three changes reversed a price; each removed a self-serve path that existed earlier in 2026. A provider chosen for Llama support in one quarter is not guaranteed to hold that support the next.
The version matrix above exists to name a fallback before that gap shows up mid-project; Best LLM API Providers covers the same access-stability question for models beyond Llama.
Compliance depth: SOC 2, HIPAA and GDPR by provider
Compliance claims split sharply once the specific certification gets checked against the marketing paragraph that surrounds it.

Fireworks AI's security documentation states SOC 2 Type II certification, HIPAA compliance, and controls mapped to GDPR and CCPA together, the only provider confirmed on all three. DeepInfra's trust center lists SOC 2 and HIPAA as compliant and GDPR as in progress.
RunPod's compliance page states it has "completed SOC 2 Type II and SOC 3 examinations, and maintains HIPAA and GDPR programs," language short of a full HIPAA claim but ahead of most GPU-rental peers.
Hyperbolic is the most transparent about a gap: its own security and compliance page states "Hyperbolic is not currently HIPAA-certified. Do not process Protected Health Information (PHI) on the platform without consulting our team first." Replicate's enterprise page uses only generic language, "built for enterprise security, privacy, and compliance," with no named certification found on a fetched page.
A regulated buyer choosing on compliance alone has one unambiguous option here: Fireworks AI. "Enterprise-ready" and "SOC 2 Type II certified" are not the same claim, and only the second one is verifiable from a fetched page.
Fine-tuning access and free-tier credits by provider
Fine-tuning access and free-tier size vary enough between providers to change a build-versus-buy decision on their own, and the table below states both side by side.
Provider | Fine-tuning or LoRA for Llama | Free tier |
|---|---|---|
Together AI | LoRA SFT $8.00, LoRA DPO $20.00 per 1M training tokens | None, $5 minimum purchase |
Fireworks AI | LoRA and full-parameter SFT/DPO by model-size band | $1 signup credit |
DeepInfra | Deploy Hugging Face LoRA adapters on supported base models | Not confirmed |
Baseten | Multi-LoRA serving confirmed; Llama-specific LoRA not confirmed | Up to $2,500 (Model APIs) or $25,000 (dedicated), Seed to Series A only |
Groq | LoRA inference only, no training, limited to | Not confirmed |
Novita AI | Not confirmed for text models | Undisclosed voucher; $10 to $500 referral credit |
Replicate | Confirmed for Llama 2 in 2023; current Llama 4 support not confirmed | Not confirmed |
RunPod | Self-managed only, no packaged product | Not confirmed |
Cerebras | Listed as an enterprise feature, moot since no Llama model is hosted | $5 credit, non-Llama models only |
Together AI's billing documentation states plainly that no free trial exists and a $5 minimum purchase is required, which means its $8.00 per 1M token LoRA rate is not testable at zero cost.
Self-hosting Llama against a managed API
API pricing for Llama hosting looks cheap at a small scale, and self-hosting is widely assumed to win once volume climbs.
The arithmetic below tests that assumption against the rates collected for this article, rather than repeating it as received wisdom.
GPU rental economics
RunPod's on-demand A100 80GB rents at $1.59 per hour and its H100 at $3.29, both confirmed on RunPod's own pricing page and both the lowest of the four providers checked. Baseten's H100 80GB, at $6.4998 per hour, is the most expensive rate found, 49% above RunPod's for the same GPU class.
Lambda sits between the two, at $2.79 for an A100 and $3.99 for an H100.

The VRAM math

VRAM here is derived, not vendor-published: the parameter counts stated above, multiplied by bytes per parameter at the stated precision, plus a 20% inference overhead, a convention Hugging Face's own Accelerate documentation attributes to EleutherAI's transformer math reference.
Llama 3.2 1B needs 0.6GB at 4-bit, fitting on a consumer GPU. Llama 3.1 405B needs 243GB, a 405 times increase for a parameter count that is also 405 times larger.
Llama 4 Scout and Maverick complicate this math in a way the active-parameter count alone hides: both are mixture-of-experts models, using only 17 billion active parameters per token, but VRAM still has to hold the full expert set, 109 billion for Scout and 400 billion for Maverick, since an unused expert still occupies memory.
Sizing Maverick by its active-parameter count alone would undersize a deployment by 23.5 times.
Does self-hosting actually win
RunPod's two-GPU setup, 2x A100 80GB at $1.59 per hour each, costs $3.18 per hour combined, or $2,289.60 across a continuous 720-hour month. CloudClusters, a GPU hosting vendor, reports a vLLM throughput range of 295.52 to 990.61 tokens per second for Llama 3.3 70B on two A100 80GB or two H100 GPUs; this article has not independently verified that figure and uses only its low end.
At 295.52 tokens per second sustained for a month, the rig generates at most 765,987,840 tokens, putting its cost at $2.99 per 1M tokens at full use. T
hat is higher than every managed price in the comparison table, including SambaNova Cloud's $1.20 output rate, the most expensive one found. Not close.
The calculation changes for workloads that need data inside a specific network boundary, need a model with no managed listing, or run below full utilization, where the fixed hourly cost makes the per-token math worse, not better.
What we did not measure
This article ran no benchmark of its own: every throughput, price and compliance figure above came from a provider's published page, not a test LLM Waves Research executed.
Uptime and service-level agreements were not compared, since no provider publishes a public uptime percentage suitable for a table.
Real-time GPU queue depth was not measured, since that state changes by the hour and any snapshot would be stale before publication.
Output quality or benchmark scores for any Llama model were not evaluated on any endpoint. Non-English performance was not assessed, a gap noted rather than filled with an estimate.
Which provider fits your situation
For developers testing quickly: DeepInfra, the lowest confirmed Llama 3.3 70B price at $0.32 per 1M output tokens, with a lineup running through Llama 4 Maverick.
For startups: DeepInfra or Novita AI, both under $0.40 per 1M output tokens. Novita AI's 12,000-token cap on that model disqualifies it for workloads built around the full 128K window.
For enterprise and regulated buyers: Fireworks AI, the only provider confirmed on SOC 2 Type II, HIPAA and mapped GDPR controls together.
For high-volume workloads: DeepInfra, not self-hosting. A two-GPU RunPod rig costs $2.99 per 1M tokens at full utilization, above every managed price found, so volume alone does not justify GPU rental unless data residency or a missing managed listing forces it.
For low latency: Groq, situationally. Its own documentation states 560 tokens per second for llama-3.1-8b-instant and 280 for llama-3.3-70b-versatile, the fastest published figures here, but both now sit behind an enterprise sales gate with no public price.
For on-device deployment: insufficient data. A hosting comparison covers off-device inference by definition; size on-device hardware from the VRAM table above instead.
For self-hosting: RunPod, the lowest confirmed A100 and H100 rental rates at $1.59 and $3.29 per hour, alongside completed SOC 2 Type II and SOC 3 examinations.
For non-English workloads: insufficient data. No provider published a non-English benchmark for any Llama model, and none was independently evaluated here.
How much VRAM do you need to host Llama 70B?
Llama 3.1 and Llama 3.3 70B need 42GB of VRAM at 4-bit precision, 84GB at 8-bit, and 168GB at 16-bit, derived from a 70 billion parameter count and a 20% inference overhead.
A single 80GB GPU covers the 4-bit case with headroom; the 16-bit case needs at least two.
Is self-hosting Llama cheaper than API pricing?
Not for Llama 3.3 70B here. A two-GPU RunPod rig running continuously costs $2.99 per 1M tokens at full utilization, calculated from RunPod's $1.59-per-hour A100 rate and a GPU vendor's published throughput figure.
Every managed Llama 3.3 70B price collected, from DeepInfra's $0.32 output rate to SambaNova Cloud's $1.20, is lower.
Does hosting Llama models carry an environmental cost?
Yes, and it scales faster than model size alone suggests. A measurement study on arXiv recorded 0.603 watt-hours for a long-prompt query against Llama 3.1 8B, 11.628 against Llama 3.1 70B, and 20.757 against Llama 3.1 405B, a 34-fold increase between the smallest and largest model at the same prompt length.
This article has not reproduced the measurement and reports it as the paper's own figure.
Which Llama hosting providers are HIPAA or SOC 2 compliant?
Fireworks AI, Together AI, DeepInfra and Baseten all state SOC 2 Type II certification and HIPAA compliance on a fetched primary-source page.
RunPod states completed SOC 2 Type II and SOC 3 examinations with HIPAA and GDPR handled as ongoing programs rather than a discrete certification.
Groq offers a signed HIPAA business associate addendum without a confirmed SOC 2 claim. Hyperbolic explicitly states it is not HIPAA-certified.
The next refresh
Every figure above is a snapshot from September 2, 2026, on pages that change without changelogs. Three providers already changed their Llama access inside the past year; a fourth could do the same before this page's next check.
The comparison worth trusting is the live provider data this article was built from, checked against your own workload rather than taken as a permanent ranking.
