Image (15): text-to-image, image-to-image, sketch-to-image, inpaint, outpaint, object removal, background replace, face swap, wardrobe/hairstyle change, style transfer, upscale, face restoration, image blend, prompt extraction, batch + seed.
Video (23): text/image-to-video, start+end frame control, keyframing, motion reference, performance transfer, trajectory control, 2D→3D camera restaging, conversational footage editing, video inpainting, character swap, background replace + relight, relight, roto/green screen, extend, lip-sync, upscale to 4K and 8K, HDR mastering, multi-shot sequences, viral templates, music video mode, podcast-to-video.
Audio (5): text-to-speech, voice cloning (consent-gated), music, sound effects, ambience.
Every tool declares the capability it needs rather than naming a provider. When no configured model has that capability, the tool tells you plainly instead of letting you spend credits on a request that would be silently approximated by a text prompt.
Without any provider API key, everything routes to a local mock adapter. The whole pipeline runs, the output is a labelled placeholder, and the UI says so throughout.