Baseten Joins Hugging Face Inference Providers
Baseten, an AI infrastructure platform, has become an official Inference Provider on the Hugging Face Hub, enhancing the serverless inference capabilities of the ecosystem. This collaboration enables developers to utilize a variety of models directly from model pages and via Hugging Face SDKs (huggingface_hub >= 1.26.1 for Python and @huggingface/inference for JavaScript). Currently, Baseten focuses on conversational and text-generation tasks, featuring well-known open-weight LLMs such as Kimi K3, DeepSeek V4 Flash, and GLM-5.2, with plans for more tasks in the future. Users can either set their own API keys or route requests through Hugging Face, incurring charges to their HF account. The integration is compatible with agent harnesses like Pi, OpenCode, Hermes Agents, and OpenClaw. PRO users receive $2 in monthly Inference credits, while free users have a limited quota. Feedback is welcomed through a discussion platform.
Key facts
- Baseten is now a supported Inference Provider on Hugging Face Hub.
- Integration covers serverless inference, training, and more.
- Initial support for conversational and text-generation tasks.
- Models include Kimi K3, DeepSeek V4 Flash, GLM-5.2.
- Available via Python and JavaScript SDKs.
- Users can set custom API keys or route through HF.
- PRO users get $2 Inference credits monthly.
- Integrates with agent harnesses like Pi, OpenCode, Hermes Agents, OpenClaw.
Entities
Institutions
- Baseten
- Hugging Face