Sign inStart creating

Using your own provider keys

Plug in your own API key or local GPU and pay an orchestration fee instead of full model cost.

Open it →

Settings → Providers. Available on Pro, Studio and Enterprise.

Add your own key for a supported provider, or point at any OpenAI- or Google-compatible endpoint — Ollama, vLLM, LM Studio, a relay, or a GPU rig in a cupboard.

Generations on your own key bypass Firefly Reels's per-model billing. We charge a 15% orchestration fee for the graph, the router, the Bibles and the continuity engine — the parts we actually provide — rather than for compute you have already bought.

If you have a provider contract, enterprise Vertex credits or your own hardware, paying twice is a reason not to use the platform at all. This removes it.

### Model discovery

A custom endpoint is asked what models it has. Modality is inferred from the model name and shown as guessed for confirmation — a listing route almost never says whether a model makes pictures or prose, and guessing silently routes an image job to a text model and fails confusingly at generation time instead of obviously at setup time. Confirm it once and a re-discovery never overwrites your answer.

Only confidently-classified generative models are enabled on discovery. An endpoint serving forty chat models should not quietly add forty routable entries.

### Key pools

Add a second key for the same provider and it becomes a pool. When one is rate-limited or exhausted the next takes over instead of the job failing. Exhausted keys rest for fifteen minutes and are then retried.

### What does not change

Moderation. Everything generated on a personal key passes exactly the same safety pipeline. Your own key changes who pays, never what is allowed.

Did this answer your question?

Related

All help articlesAsk the community