Baseten on Hugging Face Inference Providers 🔥

2026-08-20 · Hugging Face

Baseten on Hugging Face Inference Providers 🔥

Hugging Face has announced that Baseten is now an officially supported Inference Provider on the Hugging Face Hub. This addition expands the platform’s serverless inference ecosystem, allowing users to access Baseten’s models directly from model pages and through integrated SDKs.

About Baseten

Baseten is an AI infrastructure platform that provides serverless AI, training capabilities, and more. It maintains a comprehensive catalog of frontier models, making it simple for developers to integrate diverse AI functionalities into applications with minimal setup.

Supported Models and Tasks

As part of the initial launch, Baseten supports conversational and text-generation tasks on Hugging Face. This enables access to popular open-weight large language models including:

  • Kimi K3
  • Latest DeepSeek V4 Flash
  • GLM-5.2
  • And many others

Baseten natively supports a wide range of model types, from LLMs to text-to-speech. Support for additional tasks will be rolled out in the near future.

The full list of models supported by Baseten is available on their page. You can also follow Baseten on Hugging Face: https://huggingface.co/baseten.

How to Use Baseten on Hugging Face

Website UI

In your Hugging Face account settings, users can:

  • Add their own API keys for providers they have accounts with. If no custom key is set, requests are routed through Hugging Face.
  • Reorder providers by preference. This ordering affects widgets and code snippets on model pages.

Model pages display compatible third-party Inference Providers, sorted according to user preference.

There are two distinct operating modes:

1. Custom Key: Requests go directly to the inference provider using your personal API key. Billing is handled by that provider.

2. Routed by HF: Authenticate with your Hugging Face token. No provider token is needed, and charges are applied to your HF account.

Using the Client SDKs

Baseten is available through the official Hugging Face SDKs:

  • Python: `huggingface_hub` (version ≥ 1.26.1)
  • JavaScript: `@huggingface/inference`

The following examples demonstrate how to call the latest DeepSeek V4 Flash model via Baseten using an OpenAI-compatible client.

Python Example


import os
from openai import OpenAI

client = OpenAI(
    base_url="https://router.huggingface.co/v1",
    api_key=os.environ["HF_TOKEN"],
)

completion = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4-Flash-0731:baseten",
    messages=[{
        "role": "user",
        "content": "Write a Python function that returns the nth Fibonacci number using memoization."
    }]
)
print(completion.choices[0].message)

JavaScript Example


import { OpenAI } from "openai";

const client = new OpenAI({
    baseURL: "https://router.huggingface.co/v1",
    apiKey: process.env.HF_TOKEN,
});

const chatCompletion = await client.chat.completions.create({
    model: "deepseek-ai/DeepSeek-V4-Flash-0731:baseten",
    messages: [{
        role: "user",
        content: "Write a Python function that returns the nth Fibonacci number using memoization.",
    }],
});

console.log(chatCompletion.choices[0].message);

Integration with Agent Harnesses

Hugging Face Inference Providers are integrated into most major Agent Harnesses, including Pi, OpenCode, Hermes Agents, OpenClaw, and others. This allows developers to use Baseten-hosted models directly in their preferred agent frameworks without additional integration code.

Billing

  • When using a provider’s own API key, you are billed directly by that provider (e.g., your Baseten account).
  • When routing requests through Hugging Face, you pay only the standard provider API rates. Hugging Face applies no additional markup.

The platform notes that revenue-sharing agreements with provider partners may be established in the future.

PRO User Benefit: Hugging Face PRO subscribers receive $2 worth of Inference credits every month. These credits can be used across all supported providers. Free users receive a small inference quota, but upgrading to PRO is recommended for access to ZeroGPU, Spaces Dev Mode, significantly higher limits, and more.

Feedback

The Hugging Face team welcomes community feedback. Please share your thoughts in the dedicated discussion: https://huggingface.co/spaces/huggingface/HuggingDiscussions/discussions/49.

This announcement was authored by members of the Baseten team including Alex Ker, Roland Crosby, Sid Shanker, Johan, Célina Hanouti, Simon Brandeis, Lucain Pouget, Wauplin, and Merve.

(Word count: 612)

Source