API OPERATIONAL

One API.
Every AI model.

A unified AI gateway for developers who want to ship across providers without rebuilding their infrastructure.

OpenAI compatibleAnthropic Messages APIStreaming support
CURL
curl https://api.llmflux.dev/v1/chat/completions \
  -H "Authorization: Bearer $LLMFLUX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "messages": [{
      "role": "user",
      "content": "Explain quantum computing simply."
    }]
  }'
gateway readyPOST /v1/chat/completions

DESIGNED FOR THE REALITIES OF PRODUCTION AI

One keyfor supported API surfacesModel choicewithout integration sprawlUsage logsfor every metered requestFailoveracross configured upstreams

Model access

Change the model.
Not your code.

Use a single endpoint, retain a familiar request shape, and pick from the models enabled in the gateway catalogue.

MODEL CATALOGUE LIVE CONFIGURATION
OpenAI-compatibleChat CompletionsResponses
Anthropic-compatibleMessagesStreaming
Catalogued modelsProvider / familyInput / output
Automatic selectionmodel: "auto"Gateway routing
The available catalogue is maintained by your configured upstream providers.

Request architecture

One control plane.
A wider model surface.

Route through a consistent API while LLMFLUX handles the provider-facing complexity behind it.

YOUR APPLICATIONPOST /v1/chat/completions
LLMFLUX GATEWAYAuthentication · routing · metering
OpenAI compatibleModels endpoint
Anthropic compatibleMessages endpoint
Provider networkConfigured upstreams

Production primitives

Infrastructure that stays
out of your way.

Practical controls for applications that need to keep moving as models and providers change.

01

Unified API

Use OpenAI-compatible Chat Completions and Responses endpoints, plus the Anthropic Messages API.

02

Provider failover

The gateway retries eligible upstream failures so a single provider incident does not stop your application.

03

Streaming

Stream model output over the same interfaces your application already uses.

04

Usage visibility

Inspect requests, tokens, status codes, models, and providers from your account.

05

Controlled access

Create and revoke API keys without exposing provider credentials to your applications.

06

Rate protection

Plan-aware limits are enforced before a request is sent upstream.

Explore the interface

Build with a request,
not a rewrite.

The playground is a UI preview. Create an account to send authenticated requests with your own API key.

Create an account
PLAYGROUNDPREVIEW
Requests run with your API keyRun in dashboard ↗
RESPONSE

Connect through one stable interface, then select the model and routing behavior that fits the request.

Why a gateway

Less integration.
More optionality.

WITHOUT A GATEWAY
Your application├─ provider SDKs├─ credential handling├─ retry behavior├─ usage aggregation└─ model migration work
WITH LLMFLUX
Your application└─ LLMFLUX API   ├─ model catalogue   ├─ upstream routing   ├─ request logging   └─ API key controls

Pricing

Start at the endpoint.

Choose the path that fits your current stage, then grow with the infrastructure.

FLEXIBLE USAGE

Usage

For applications in production.

Pay as you gopooled model accessGet an API key →

Enterprise

For teams with operational requirements.

Customtalk to the teamContact us →

FAQ

Answers for
builders.

For implementation details, see the API reference.

Open API docs
What is an AI API gateway?

An AI API gateway gives an application one interface for accessing AI models while centralizing routing, credentials, usage tracking, and reliability behavior.

Which API formats are supported?

LLMFLUX supports OpenAI-compatible Chat Completions and Responses endpoints, as well as the Anthropic Messages API.

Can I switch models without changing my integration?

Yes. Once your application is pointed at LLMFLUX, select a different model in the request or use automatic model selection where available.

Do you support streaming?

Yes. Streaming requests are proxied through the gateway for the supported text-generation endpoints.

Where can I inspect usage?

The dashboard includes usage summaries and request logs so you can review tokens, latency-related request details, and outcomes.

Ready when you are

Build once.
Ship with every model.

Give your application a durable path into the AI ecosystem without coupling it to every provider.