v2.0 / SELF-HOSTED AI GATEWAY / MIT

One endpoint,
every channel
and credential

Thirty-two channels, dozens of credentials, and four conversation protocols sit behind one base URL. The gateway handles scheduling, retries, cooldown, and usage accounting; applications only configure the base URL and AccessKey.

Your application1 BASE URL1 AccessKeyAUTHGPT-LOADSCHEDULE / RETRY / COOLDOWNAFFINITY / LOG / USAGEOfficial APIs4 CHANNELSCloud platforms3 CHANNELSModel services19 CHANNELSSubscriptions4 CHANNELSFIG. 1 — ROUTINGROUTED → Official APIs0 REQ
35
built-in channels
4
conversation protocols
1
Go binary
2
connection settings
SQLite / MySQL / PostgreSQLCredentials encrypted locally · Requests go to your chosen upstreams
01Support

Supported by these partners

GPT-Load is open source under the MIT License. These sponsors help cover infrastructure, model credits, and development time.

Full strength, stable, nothing watered down. Global SOTA models at a lower cost.

Principal Partner

Modelflare supports GPT, Claude, Gemini, Grok, and other leading models through common compatible protocols, plus 40+ image and video models. Cache compensation is a highlight: selected OpenAI groups (price-first, stable, and premium) guarantee daily cache hit rates of 65%, 75%, or 85%, and if eligible requests fall short, Modelflare refunds the difference, visible on the usage dashboard after settlement. Special offers: a single top-up of US$20 unlocks the GPT discount group at a 0.015x price multiplier, and a single top-up of US$50 unlocks the Claude discount group at a 0.03x multiplier.

Register via this exclusive link to get started →

Text, image, and video AI in one platform

OfoxAI is a unified AI API platform bringing together text, image, and video models from multiple providers. With OpenAI-compatible endpoints and native Anthropic and Gemini interfaces, developers can access models for AI applications, agents, and content creation through one platform.

Explore OfoxAI models and APIs →

PackyCode

Access leading AI models through PackyCode with one API endpoint and one API key. Enjoy fast, reliable access with automatic failover and dedicated high-speed routes for Codex and Claude Code. Get started with $1 in free credits, a discount on your first top-up, and savings of up to 80% on eligible routes. Pay in RMB with no currency conversion markups or extra top-up fees.

Sign up and start building →

APIMart

Low-cost APIs for AI image and video generation, with GPT-Image-2 from $0.006 per image. One asynchronous API handles both images and video, scales to large batches, and has no monthly fee.

Learn more →
02Capabilities

Six jobs the gateway handles for you

Each application would otherwise implement these jobs separately—or skip them until something breaks. Put them in the gateway once and every client benefits.

01

One scheduler, two credential types

API keys and subscription accounts for Codex, Claude, Antigravity, and Grok share one pool and the same scheduling, retry, cooldown, and health-isolation behavior.

02

Clients keep their protocols

OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Gemini are supported. The gateway uses a native or converted route as declared by the selected channel. Existing clients usually need only a new base URL and AccessKey.

03

Failures route themselves away

Weighted rotation, session affinity, retries, cooldown, and automatic blacklisting isolate a limited or invalid credential before it can affect the full route.

04

Every call stays observable

Inspect health, routes, request logs, usage summaries, and per-model cost estimates down to the credential handling a request.

05

One binary, data under your control

The management UI ships inside the Go binary. Start with SQLite or connect MySQL or PostgreSQL; upstream credentials are encrypted locally and never pass through a third party.

06

35 built-in channels

Official APIs, cloud platforms, model services, subscription accounts, and gateway channels including GPT-Load, New API, CLIProxyAPI, Sub2API, and OpenAI Compatible are built in. See the support matrix for exact capabilities.

03Channels & protocols

Four conversation protocols in, 35 channels out

The gateway connects client protocols to upstream channels. Protocol and operation support differs by channel; the capability matrix is the exact reference.

Client protocols
OpenAI Chat Completions
POST /v1/chat/completions
The most common endpoint for compatible clients
OpenAI Responses
/v1/responses/**
Namespace-preserving pass-through with stateful resources
Anthropic Messages
POST /v1/messages
Native entry point for Claude Code and similar clients
Gemini
/v1beta/models/…
Includes generateContent and streaming variants
Upstream channels
Official APIs4
  • OpenAI
  • Anthropic
  • Gemini
  • xAI
Cloud platforms3
  • Azure OpenAI
  • AWS Bedrock
  • Google Vertex AI
Model services19
  • DeepSeek
  • Moonshot AI
  • SiliconFlow
  • Zhipu AI
  • Alibaba Cloud
  • Volcengine
  • OpenRouter
  • Groq
  • Jev
  • Cerebras
  • Mistral
  • Nebius
  • Parasail
  • Wafer
  • Hugging Face
  • Cohere
  • Cline
  • OpenCode Go
  • OpenCode Zen
Subscriptions4
  • Codex
  • Claude
  • Antigravity
  • Grok
+ Gateway and compatible channels — GPT-Load · New API · CLIProxyAPI · Sub2API · OpenAI CompatibleView the capability matrix →
04Setup

Configure it in three steps

There are only two things to manage: Groups face upstream, while AccessKeys face applications. Scheduling, retries, and usage accounting stay between them.

01

Create a Group

Choose an upstream channel and add its API keys. Subscription accounts such as Codex and Claude use OAuth, then join the same scheduler.

Group → channel + credential pool
02

Choose exposed models

Select the models this Group exposes, discover them from upstream, and adjust weight, timeouts, retries, and other runtime policy when needed.

Group → models + policy
03

Issue an AccessKey

Set the Groups and client protocols it may use, add rate or cost limits, and give the generated key to your application. That is all the application needs to know.

AccessKey → authorization + limits
05Get started

One command to start, two client settings to change

No external database and no separate frontend deployment are required. The management UI is embedded in the same binary.

① Start the service
Docker Compose
git clone --depth 1 --branch main \
  https://github.com/tbphp/gpt-load.git
cd gpt-load && cp .env.example .env
docker compose up -d

# Read the generated management key
docker compose exec gpt-load \
  sh -c 'cat /app/data/auth.key'
② Connect a client
Python — OpenAI SDK
client = OpenAI(
    base_url="http://127.0.0.1:3001/v1",   # change this
    api_key="sk-gl-••••a5df",               # change this
)

# Anthropic clients use /v1/messages
# Gemini clients use /v1beta/models/…

Keep each client's normal authentication style: Authorization: Bearer · x-api-key · x-goog-api-key · Gemini key query parameter.

Native binaries

Five build targets

Linux and macOS on both amd64 and arm64, plus Windows on amd64. Verify the bundled SHA256SUMS before running a downloaded binary; Windows also has an installer that registers a service.

Database

Start on SQLite, switch when needed

Leave DATABASE_DSN empty for managed SQLite, or set a DSN for external SQLite, MySQL, or PostgreSQL.

Backup

Keep the key with the database

encryption.key decrypts channel credentials. If it is lost or replaced, encrypted credentials cannot be recovered; this release does not support master-key rotation.

Full setup and additional client examplesQuick start →
06Management UI

Configuration and observability in one place

The management UI ships with the binary. Open a browser to configure channels, inspect health and logs, and review usage and cost estimates.

FIG. 2 — Groups and channelsChannels · models · credential health

The Groups list brings channel, connection type, model count, and credential health into one view. API keys and subscription accounts share the same availability and scheduling view.

FIG. 3 — Runtime healthAvailability · cooldowns · log collection

See available, cooling-down, and blacklisted credentials at a glance, with recovery times, request-log collection, and Group health on the same page.

07Boundaries

Read these three points before production

Cost

Usage and cost are estimates

Figures are derived from upstream responses for operations and capacity planning. They are not provider invoices, cannot be used for financial reconciliation, and price changes do not recalculate history.

Scale

Designed for a single application instance

Version 2.0 guarantees correctness within one instance. Instances do not share state and horizontal scaling is not supported; split larger workloads into independent deployments.

Upgrade

1.x cannot be upgraded in place

Version 2.0 is a full rewrite and cannot open, import, or migrate 1.x data. Deploy it with a separate database, DATA_DIR, port, and volume, then switch traffic after validation.

Network boundary

The service listens on 127.0.0.1 by default and is not exposed publicly. For remote access, use a controlled network or TLS reverse proxy with appropriate ACL and firewall rules. Protect AUTH_KEY and ENCRYPTION_KEY; never put them in repositories, logs, screenshots, or public issues.

GPT-Load — Self-hosted AI Gateway