video_generate tool call. Built-in providers (xAI, FAL, OpenRouter, DeepInfra) ship as plugins. Add a new one, or override a bundled one, by dropping a directory into plugins/video_gen/<name>/.
The unified surface (one tool, two modalities)
Thevideo_generate tool exposes two modalities through one parameter:
- Text-to-video — call with
promptonly. The provider routes to its text-to-video endpoint. - Image-to-video — call with
prompt+image_url. The provider routes to its image-to-video endpoint.
How discovery works
Mibyan scans for video-gen backends in three places:- Bundled —
<repo>/plugins/video_gen/<name>/(auto-loaded withkind: backend) - User —
~/.mibyan/plugins/video_gen/<name>/(opt-in viaplugins.enabled) - Pip — packages declaring a
mibyan_agent.pluginsentry point
register(ctx) function calls ctx.register_video_gen_provider(...). The active provider is picked by video_gen.provider in config.yaml; mibyan tools → Video Generation walks users through selection. Unlike image_generate, there is no in-tree legacy backend — every provider is a plugin.
Directory structure
The VideoGenProvider ABC
Subclassagent.video_gen_provider.VideoGenProvider. Required: name property and generate() method.
The plugin manifest
The video_generate schema
The tool exposes one schema across every backend. Providers ignore parameters they don’t support.
There is deliberately no
model parameter: the backend and model are user configuration (video_gen.provider / video_gen.model), never an agent choice. Your generate() still receives model= — it is the configured model, resolved by the tool layer.
The provider’s capabilities() advertises which of these are honored. The agent sees the active backend’s capabilities in the tool description, dynamically rebuilt when the user changes backend via mibyan tools.
Model families and endpoint routing (the FAL pattern)
When your backend has multiple endpoints per “model” — like FAL, where every family (Veo 3.1, Pixverse v6, Kling O3) has both a/text-to-video and an /image-to-video URL — represent each family as one catalog entry. Your generate() picks the right endpoint based on whether image_url was passed:
veo3.1 once in mibyan tools. The agent never thinks about endpoints — it just passes (or doesn’t pass) image_url.
Selection precedence
For per-instance model knobs (seeplugins/video_gen/fal/__init__.py):
<PROVIDER>_VIDEO_MODELenv varvideo_gen.<provider>.modelinconfig.yamlvideo_gen.modelinconfig.yaml(when it’s one of your IDs)- Provider’s
default_model()
model= keyword your generate() receives is the outcome of this resolution — a model in the agent’s tool call is ignored, so the LLM cannot switch backends or billing tiers on its own.
Response shape
success_response() and error_response() produce the dict shape every backend returns. Use them — don’t hand-roll the dict.
Success keys: success, video (URL or absolute path), model, prompt, modality ("text" or "image"), aspect_ratio, duration, provider, plus extra.
Error keys: success, video (None), error, error_type, model, prompt, aspect_ratio, provider.
Where to save artifacts
If your backend returns base64, usesave_b64_video() to write under $mibyan_HOME/cache/videos/. For raw bytes from a follow-up HTTP fetch, use save_bytes_video(). Otherwise return the upstream URL directly — the gateway resolves remote URLs on delivery.
Testing
Drop a smoke test undertests/plugins/video_gen/test_<name>_plugin.py. The xAI and FAL tests show the pattern — register, verify catalog, exercise routing both with and without image_url, assert clean error responses on missing auth.
