Open SourceInference11k
OpenLLM
Run any open LLM as an OpenAI-compatible API endpoint, anywhere.
OpenLLM, by the BentoML team, lets you run open-source and fine-tuned LLMs as OpenAI-compatible API endpoints with a single command. It supports popular models out of the box, optimized inference backends like vLLM, built-in chat UI, and one-click deployment to Docker, Kubernetes, or BentoCloud. It is designed to make self-hosting production LLM APIs simple and portable.
Repository
bentoml/OpenLLMLanguage
PythonWhat you'd build with it
- Spinning up an OpenAI-compatible endpoint for any supported open model with one command
- Packaging an LLM service into a Bento for reproducible Docker/Kubernetes deployment
- Serving fine-tuned models with vLLM-backed inference and a built-in chat UI
Tags
servingopenai-compatibledeploymentself-hosted