FLUX 3 video over an async HTTP API: requests, polling, keyframes, and real pricing
Status check: this is GA now, not early access BFL closed out the July Early Access phase. The first text-to-video and image-to-video release went generally available through the BFL API and selected partners on August
Status check: this is GA now, not early access
BFL closed out the July Early Access phase. The first text-to-video and image-to-video release went generally available through the BFL API and selected partners on August 4, 2026. The production model ID flux-3 appeared in the gateway's Video API format on August 13, 2026. I verified availability, field names and prices on September 24, 2026.
Those facts move, so treat the BFL release announcement and the live model page as the source of truth instead of this post. Early coverage of the July application-based rollout is still useful for launch history, but it does not describe a callable endpoint. What follows is the operational part: create, poll, download, and the schema traps in between.
What the model does
FLUX 3 is BFL's multimodal family covering video, audio, images, and action-related prediction. The video release handles text-to-video, image-to-video, and video continuation behind one provider-native endpoint.
The parts that matter for integration work:
- Up to 20 seconds at 24 fps, HD or FHD
- Synchronized audio, on by default
- Multilingual speech with lip-sync
- Multiple shots inside a single generation
- Up to ten pinned keyframes for image-to-video control
Specs
| Spec | BFL native | Gateway |
|---|---|---|
| Workflows | T2V, I2V, video continuation | T2V and I2V listed |
| Max duration | 5 to 20 s T2V/I2V; 5 to 15 s V2V | check live quickstart values |
| Frame rate | 24 fps | provider output |
| Resolution | rHD native, FHD via video upsampler; no 4K/UHD listed for FLUX 3 Video | 720p and 1080p pricing |
| Native audio | yes, default on | follows the active integration |
| Image control | 1 to 10 keyframes | verify reference-image mapping |
| Aspect ratios | 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16 | quickstart uses explicit sizes like 1280x720 |
| Invocation | async | create, poll, download |
| Model ID | n/a | flux-3 |
One clarification that keeps circulating: the "up to 4MP" spec belongs to the FLUX.2 image models. FLUX 3 Video does not list 4K/UHD output.
The three-call lifecycle
| Operation | Request | Returns |
|---|---|---|
| Create | POST https://api.cometapi.com/v1/videos |
task ID + queued status |
| Poll | GET https://api.cometapi.com/v1/videos/{task_id} |
status and progress |
| Download | GET https://api.cometapi.com/v1/videos/{task_id}/content |
the MP4 |
Set the key in a backend env var. macOS/Linux:
export COMETAPI_KEY="your_api_key"
Windows PowerShell:
$env:COMETAPI_KEY="your_api_key"
Create a five-second 720p clip with a multipart request:
curl https://api.cometapi.com/v1/videos \
-H "Authorization: Bearer $COMETAPI_KEY" \
-F "model=flux-3" \
-F "prompt=A paper boat glides across a still pond in soft morning light" \
-F "seconds=5" \
-F "size=1280x720"
The response starts a job, nothing more:
{
"id": "video_task_id",
"status": "queued"
}
Persist that ID next to your own job or user record before you poll anything. A worker restart should never cost a user a second billed generation. Then poll:
curl https://api.cometapi.com/v1/videos/{task_id} \
-H "Authorization: Bearer $COMETAPI_KEY"
Ten seconds is a sane starting interval. Terminal success states I've seen: completed, succeeded, success. Terminal failures: failed, failure, cancelled, canceled.
Download, then move the file into your own bucket rather than pinning a provider URL as the permanent asset:
curl https://api.cometapi.com/v1/videos/{task_id}/content \
-H "Authorization: Bearer $COMETAPI_KEY" \
--output flux3_output.mp4
Complete Python job runner
This creates, persists, polls, checks failure states, validates the MP4 signature, and writes to disk.
import os
import time
from pathlib import Path
import requests
api_key = os.environ["COMETAPI_KEY"]
base_url = "https://api.cometapi.com"
headers = {"Authorization": f"Bearer {api_key}"}
response = requests.post(
f"{base_url}/v1/videos",
headers=headers,
files={
"model": (None, "flux-3"),
"prompt": (
None,
"A product bottle rotates slowly on wet black stone, "
"soft rim lighting, macro lens, realistic reflections.",
),
"seconds": (None, "5"),
"size": (None, "1280x720"),
},
timeout=120,
)
response.raise_for_status()
task = response.json()
data = task.get("data") or {}
task_id = (
task.get("id")
or task.get("task_id")
or data.get("id")
or data.get("task_id")
)
if not task_id:
raise RuntimeError(f"Create response has no task ID: {task}")
while True:
response = requests.get(
f"{base_url}/v1/videos/{task_id}",
headers=headers,
timeout=60,
)
response.raise_for_status()
task = response.json()
data = task.get("data") or {}
status = str(task.get("status") or data.get("status") or "").lower()
progress = task.get("progress") or data.get("progress") or "unknown"
print(f"Status: {status or 'unknown'}; progress: {progress}")
if status in {"failed", "failure", "cancelled", "canceled"}:
raise RuntimeError(f"Video generation failed: {task}")
if status in {"completed", "succeeded", "success"} or progress == "100%":
break
time.sleep(10)
response = requests.get(
f"{base_url}/v1/videos/{task_id}/content",
headers=headers,
timeout=300,
)
response.raise_for_status()
video = response.content
if len(video) < 12 or video[4:8] != b"ftyp":
raise RuntimeError("Content response is not a non-empty MP4 file")
output_dir = Path("output")
output_dir.mkdir(parents=True, exist_ok=True)
output_path = output_dir / f"{task_id}.mp4"
output_path.write_bytes(video)
print(f"Saved: {output_path} ({len(video)} bytes)")
Image-to-video and keyframes: two schemas, not one
I2V is listed as supported through the gateway, but the current public sample only demonstrates text-to-video. Do not assume a reference-image field copied from another model's docs will be accepted unchanged. Check the live gateway reference before you ship that path.
BFL's native API is explicit: I2V uses mode i2v plus the keyframes field. One image pins the start frame, two pin start and end, and up to ten timed images storyboard a continuous clip.
curl -X POST https://api.bfl.ai/v1/flux-3-video \
-H "x-key: $BFL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "i2v",
"prompt": "They sprint through the lantern-lit alley as the camera tracks behind them.",
"keyframes": [
[0, "data:image/png;base64,"],
[8, "data:image/png;base64,"]
],
"duration": 8
}'
Keep native and gateway parameters in separate adapters. Mixing them is the most common failure I see in review. Native fields are mode, keyframes, start_video, resolution, draft; the verified gateway sample uses model, prompt, seconds, size.
Parameter reference
| Parameter | API | Controls | Guidance |
|---|---|---|---|
model |
gateway | model selection | flux-3 |
prompt |
both | scene, action, camera, audio | describe change over time |
seconds |
gateway | clip length | start at 5 to 8 s while tuning |
size |
gateway | dimensions | 1280x720 for cheap iteration |
mode |
BFL native | t2v / i2v / v2v / draft_enhance | do not send unless mapped |
duration |
BFL native | 5 to 20 s T2V/I2V, 5 to 15 s V2V |
auto supported |
resolution |
BFL native |
hd or fhd
|
FHD finishes through the upsampler |
generate_audio |
BFL native | synchronized audio | defaults to true |
draft |
BFL native | fast preview | lower-cost creative iteration |
Prompting: write a shot list, not a mood
BFL's video prompting guide asks for clear direction on subject and action, camera, scene and atmosphere, motion quality, and continuity. Audio-led scenes need dialogue, voice, sound effects, and ambience spelled out.
The structure I reuse:
Subject + Environment + Action + Camera + Lighting
+ Dialogue/Voice + Sound Effects + Ambience + Constraints
Cinematic:
A lone cyclist rides through a rain-soaked neon street at midnight.
The camera begins low beside the rear wheel, then rises into a smooth tracking shot.
Reflections stretch across wet asphalt under moving cyan and magenta light.
Audio: steady rainfall, chain noise, distant traffic, no music, no dialogue.
Keep the same rider, bicycle, jacket, and weather throughout the shot.
Product:
A premium stainless-steel espresso machine stands on a dark stone counter.
Begin with a macro close-up of water droplets on the metal housing.
Orbit clockwise as the machine brews; steam catches warm side light.
Finish on a clean three-quarter hero angle with the cup in the foreground.
Audio: pump vibration, steam hiss, ceramic contact, quiet cafe ambience.
Do not change the product shape, logo placement, material, or color.
Dialogue with native audio:
A young chef works alone in a compact Tokyo ramen shop at night.
Start close on boiling broth, then pull back as the chef sets down a bowl.
Warm tungsten lighting, natural reflections, documentary handheld motion.
The chef quietly says in Japanese: γγεΎ
γγγγΎγγγγ
Audio: bubbling broth, soft rain outside, distant street traffic.
No subtitles and no background music.
"Make a cinematic ramen shop video" leaves motion, framing, sound and continuity undefined. You cannot debug that.
Pricing
BFL prices by workflow. T2V and I2V full renders run $0.17/s HD and $0.29/s FHD, with HD Draft Mode at $0.06/s. Video continuation is $0.43/s HD and $0.54/s FHD, HD drafts at $0.12/s. The gateway lists flux-3 at $0.136/s for 720p and $0.232/s for 1080p. Re-check live pricing before any large batch.
| Workflow | HD/720p | FHD/1080p | Draft | 5 s full | 10 s full |
|---|---|---|---|---|---|
| BFL T2V | $0.17/s | $0.29/s | $0.06/s (HD) | $0.85 / $1.45 | $1.70 / $2.90 |
| BFL I2V | $0.17/s | $0.29/s | $0.06/s (HD) | $0.85 / $1.45 | $1.70 / $2.90 |
| BFL V2V continuation | $0.43/s | $0.54/s | $0.12/s (HD) | $2.15 / $2.70 | $4.30 / $5.40 |
Gateway flux-3
|
$0.136/s | $0.232/s | not listed | $0.68 / $1.16 | $1.36 / $2.32 |
In the last two columns the first figure is HD/720p and the second is FHD/1080p.
Cutting iteration cost
- Prototype at 720p, then promote only a selected prompt to 1080p.
- Five-second clips are enough to validate composition, motion and prompt interpretation.
- Change one major variable per run.
- On the native API, use Draft Mode before a full-quality render.
- Store the winning prompt and reference decisions in your own metadata.
FLUX 3 vs Wan 3.0 vs Seedance 2.5
Pick by workflow, not by a leaderboard. This is also where a unified endpoint earns its keep: keeping flux-3, Wan 3.0 and Seedance 2.5 behind one job pipeline (I route through CometAPI for exactly that) means a model swap is a config change rather than a rewrite.
| Dimension | FLUX 3 | Wan 3.0 | Seedance 2.5 |
|---|---|---|---|
| Max clip | up to 20 s T2V/I2V | up to 30 s | up to 30 s |
| Synchronized audio | yes | yes | yes |
| T2V / I2V | yes / yes | yes / yes | yes / yes |
| Reference strategy | up to 10 native keyframes | broad multimodal, Omni-Reference | large multimodal reference capacity |
| Continuation/editing | native BFL v2v | long-form and editing | extension and editing |
| Strength | motion logic, multi-shot, synced A/V | input breadth, 30 s generation | long storytelling, reference-heavy control |
| Gateway starting price | $0.136/s | $0.04/s | $0.0824/s |
| Fits | cinematic or realistic A/V shots | all-in-one multimodal pipelines | longer, identity-controlled storytelling |
Starting prices are not apples-to-apples across resolutions, so budget from each live model page's resolution table.
Production notes
Persist async jobs. Save the task ID at submission time. Worker retries must not re-bill work.
Do not poll tightly. Around ten seconds unless the live docs say otherwise. One-second polling adds load without improving perceived latency.
Validate the download. Check length and the MP4 ftyp signature before marking an asset ready. A 200 does not prove you got a video.
Separate native and gateway schemas. Distinct adapters stop native-only fields from leaking into gateway calls.
Log the full failure context. HTTP status, response body, task ID, model ID, prompt version, size, duration, internal job ID. Redact the key.
Run a fixed evaluation set. 10 to 30 prompts covering camera motion, people, products, typography, dialogue, high-motion scenes and required aspect ratios. Re-run it on every model or integration change, and repeat important prompts, because generation is stochastic.
Vendor numbers worth knowing
BFL reports a text-to-video Elo of 1135 in its all-vs-all human-preference evaluation, and a tie with Seedance 2.0 on image-to-video preference. Useful positioning, but vendor-run preference testing measures perceived output quality, not queue reliability, gateway latency, cost consistency, or generation-to-generation stability. Your own prompt set is the only benchmark that predicts your production.
FAQ
Model ID? flux-3.
Endpoints? POST /v1/videos, then GET /v1/videos/{task_id} and GET /v1/videos/{task_id}/content.
Synchronous? No. Submit, persist, poll, download.
Max duration? 5 to 20 s for T2V/I2V, 5 to 15 s for continuation.
Audio? Yes, synchronized and enabled by default in the native API.
Can I send BFL keyframes through the gateway unchanged? Do not assume it. The schemas differ. Verify the current quickstart first.
Cost of five seconds? $0.68 at 720p or $1.16 at 1080p at current gateway rates.
What I would ship first
Short 720p text-to-video request, durable task ID, conservative polling, download with signature validation, then prompt templates, object storage, retry logic and a repeatable evaluation set. Add keyframes and continuation once I2V field mapping is confirmed against the live reference.
Originally published at cometapi.com
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.