Settings → Providers. Available on Pro, Studio and Enterprise.
Add your own key for a supported provider, or point at any OpenAI- or Google-compatible endpoint — Ollama, vLLM, LM Studio, a relay, or a GPU rig in a cupboard.
Generations on your own key bypass Firefly Reels's per-model billing. We charge a 15% orchestration fee for the graph, the router, the Bibles and the continuity engine — the parts we actually provide — rather than for compute you have already bought.
If you have a provider contract, enterprise Vertex credits or your own hardware, paying twice is a reason not to use the platform at all. This removes it.
### Model discovery
A custom endpoint is asked what models it has. Modality is inferred from the model name and shown as guessed for confirmation — a listing route almost never says whether a model makes pictures or prose, and guessing silently routes an image job to a text model and fails confusingly at generation time instead of obviously at setup time. Confirm it once and a re-discovery never overwrites your answer.
Only confidently-classified generative models are enabled on discovery. An endpoint serving forty chat models should not quietly add forty routable entries.
### Key pools
Add a second key for the same provider and it becomes a pool. When one is rate-limited or exhausted the next takes over instead of the job failing. Exhausted keys rest for fifteen minutes and are then retried.
### What does not change
Moderation. Everything generated on a personal key passes exactly the same safety pipeline. Your own key changes who pays, never what is allowed.