The API gateway for serving LLM inference requests with OpenAI-compatible HTTP and KServe gRPC endpoints.
See docs/components/frontend/ for documentation.
The API gateway for serving LLM inference requests with OpenAI-compatible HTTP and KServe gRPC endpoints.
See docs/components/frontend/ for documentation.