All open source
Open SourceInference11k

OpenLLM

Run any open LLM as an OpenAI-compatible API endpoint, anywhere.

OpenLLM, by the BentoML team, lets you run open-source and fine-tuned LLMs as OpenAI-compatible API endpoints with a single command. It supports popular models out of the box, optimized inference backends like vLLM, built-in chat UI, and one-click deployment to Docker, Kubernetes, or BentoCloud. It is designed to make self-hosting production LLM APIs simple and portable.

Repository

bentoml/OpenLLM

Language

Python

What you'd build with it

  • Spinning up an OpenAI-compatible endpoint for any supported open model with one command
  • Packaging an LLM service into a Bento for reproducible Docker/Kubernetes deployment
  • Serving fine-tuned models with vLLM-backed inference and a built-in chat UI

Tags

servingopenai-compatibledeploymentself-hosted