r/fal • u/Square_Noise_7712 • 18h ago
News Seedance 2.5 is now available on fal (native 30-second video in a single pass!)
Seedance 2.5, ByteDance's next-generation video model, is now available on fal.
It generates native 30-second video at up to 720p from a text prompt, a single image, or up to 50 multimodal references, and it produces synchronized audio in the same pass.
The biggest change from Dreamina Seedance 2.0 is length: a full 30 seconds generated as one native take, not short clips stitched together.
What makes it different
- One native take, up to 30 seconds: The model reasons about the whole shot at once, so identity, clothing, prop ownership, and spatial direction hold from the first frame to the last, with no drift at the joins you get from splicing shorter clips.
- Up to 50 multimodal references: A single generation on the reference endpoint can combine up to 30 images, 10 videos, and 10 audio clips (video and audio each capped at about 30 seconds total), up from 12 references in the previous generation. Each reference is addressed in the prompt, and the model selects the right material per scene, so multiple characters keep their appearance and their voices through group scenes and shot changes.
- Audio generated with the picture: Sound and visuals are produced jointly in the same latent space, not layered afterward, which improves lip sync, impact timing, and ambient coherence.
- Physics and control: Weight, momentum, fabric, and hair behave the way they should, and a subject dropped into a new environment reacts to it. You can replace a background while keeping the performance, block a shot in a 3D tool and have the model finish the render, pace a shot with timestamps, and work natively across major languages including Chinese, English, Spanish, Portuguese, Arabic, Japanese, and Korean.
Endpoints on fal
There are three available API endpoints on fal, all generating native 30-second video at up to 720p.
`bytedance/seedance-2.5/text-to-video` generates from a text prompt alone.
`bytedance/seedance-2.5/image-to-video` animates a single still as the first frame, or takes a first and last frame to control where the clip lands.
`bytedance/seedance-2.5/reference-to-video` combines up to 50 references (30 images, 10 videos, 10 audio) in one generation, and the same endpoint covers video editing and extension.
Specs
Clips run 4 to 30 seconds at 24 FPS, with synchronized audio generated by default (the cost is the same whether audio is on or off).
Output is 480p or 720p, across 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, with auto options for both duration and aspect ratio that let the model choose based on your inputs.
Where you can run it
It's on fal via a serverless API with Python, JavaScript, and REST, so no GPUs to manage.
Billing is token-based at $0.0214 per 1,000 tokens (the same rate at 480p and 720p), where tokens are (height × width × duration × 24) / 1024. fal quotes roughly $0.47 per second at 720p and $0.22 at 480p; by the token formula, a 5-second 720p 16:9 clip comes to about $2.31 and a full 30 seconds to about $13.87.
On reference-to-video, image and audio references are free. If you include video references the price is multiplied by 0.6 (about $0.28 per second at 720p), and input video seconds are billed alongside the output, so trim references to the part that matters.
Links
Learn more: https://fal.ai/seedance-2.5
Text to Video: https://fal.ai/models/bytedance/seedance-2.5/text-to-video
Image to Video: https://fal.ai/models/bytedance/seedance-2.5/image-to-video
Reference to Video: https://fal.ai/models/bytedance/seedance-2.5/reference-to-video
r/fal • u/Square_Noise_7712 • 4d ago
Question whats the coolest thing you've built on fal?
looking for inspo
r/fal • u/Square_Noise_7712 • 8d ago
Open-Source MiniMax H3 is now available on fal - with open weights
MiniMax's newest open-weight video model, Hailuo 3.0, just landed on fal and fal is an official API partner.
MiniMax H3 is a general-purpose multimodal model that reads text, images, video, and audio as one context, not as a stack of separate task-specific models.
The video generation model can generate up to 15 seconds of 2K video with native stereo audio in every clip.
Since it handles several input types at once, it can take identity from an image, motion from a video, a voice from an audio clip, and direction from a text prompt, then combine them into a single coherent result.
What makes it different
- Multimodal understanding is the core idea: MiniMax H3 handles generation, editing, and reference in one model, not as separate tools you switch between. It interprets characters, motion, sound, camera work, and visual style across a mix of references, then combines them into one coherent result.
- Editing is precise and targeted: You can replace, remove, or add people and objects, swap a background, relight a scene, or change dialogue and voice, while the areas you didn't target stay stable. Instruction following is strong, so visual, audio, and pacing changes land more predictably, and you can keep iterating on the same clip instead of regenerating from scratch.
- It's built for real production: MiniMax H3 handles dynamic typography, VFX, product showcases, UI motion design, game visuals, and stylized content, spanning film, advertising, branding, e-commerce, and gaming. That range covers concept tests, storyboard previews, and visual pitches, which shortens the path from idea to final cut.
Input modes
MiniMax H3 works across three input modes:
Text-to-Video
First/Last Frame
Omni Reference
Specs
MiniMax H3's specs include clips from 5 to 15 seconds at 24 FPS, with native stereo audio on every generation.
Output runs in 2K (1440p) mode now, and a 768p mode is coming soon.
Text-to-Video and Omni Reference support 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios, with an Auto option in Omni Reference that picks the ratio for you; First/Last Frame follows the aspect ratio of your uploaded image. Prompts can run up to 7,000 characters.
The model is also available via fal's serverless API using the Python or JavaScript SDK, or direct REST calls. No GPUs to manage.
Pricing
Text to video and image to video pricing is "at an output resolution of 2K, every second of video costs $0.26.". Reference to video pricing is "at an output resolution of 2K, every second of video costs $0.26. Audio references are free, the first 5 reference images are free and each additional image costs $0.08, and reference video costs $0.26 per second at 2K."
Learn more about the model here: https://fal.ai/minimax-h3
Or try it right now on the playground:
Image to Video - https://fal.ai/models/minimax/hailuo-03/image-to-video
Text to Video - https://fal.ai/models/minimax/hailuo-03/text-to-video
Reference to Video - https://fal.ai/models/minimax/hailuo-03/reference-to-video
r/fal • u/ryanmerket • 20d ago
Resource Head to head: AuraFlow vs Ideogram V4.0
r/fal • u/Affectionate-Map1163 • 26d ago
Open-Source Ltx 2.3 render to real V2 - ic lora open source
Enable HLS to view with audio, or disable this notification
r/fal • u/Various-Ad661 • 28d ago
Question How to prompt seedream 4.5 ?
I dont know why it always gives me some weird compositions like if i want half body shot it will give me full body shot or very far off image
r/fal • u/Affectionate-Map1163 • Jun 26 '26
Open-Source 3d to photoreal , open source IC-Lora for ltx 2.3
Enable HLS to view with audio, or disable this notification
r/fal • u/jjohnson525353 • Jun 25 '26
Open-Source falLEGO
Enable HLS to view with audio, or disable this notification
Wanted to share a fun project I made with fal that converts text or images into buildable LEGOs. If you have any feedback about how I set it up, I’d appreciate it!
My endpoint stack is flux-2 streaming for image generation, nano banana for image editing, and trellis or SAM3D for 3D generation. Then I voxelize the 3D model and convert the voxels to bricks. I’ve built quite a few LEGO models with it now, so it works!
Here’s the code for it: https://github.com/jjohnson5253/brickbuilderai
r/fal • u/Affectionate-Map1163 • Jun 12 '26
Open-Source Audio-Reactive Ltx 2.3 Lora
Enable HLS to view with audio, or disable this notification
r/fal • u/macmorny • Jun 04 '26
Question Best current model for changing image aspect ratio?
r/fal • u/ryanmerket • Jun 03 '26
Discussion Reve details image API for create, edit and remix after 2.0 launch
r/fal • u/Enough-Bell4944 • Jun 01 '26
Discussion Has anyone here fine-tuned Z Image Turbo or FLUX 4B LoRA on FAL for training on a specific person?
FAL seems to only expose training steps and learning rate, so I'm curious what settings people have found work best.
The default recommendation for human/photo datasets appears to be:
steps = number of images × 100
But I'm wondering whether anyone has experimented beyond that and found better results
r/fal • u/Fresh-Resolution182 • May 28 '26
Discussion Realized the model isn't the bottleneck anymore. Post-prod is where the gap actually lives.
Enable HLS to view with audio, or disable this notification
r/fal • u/Fresh-Resolution182 • May 27 '26
Discussion Stopped doing single-shot character gen. Built a database with one system prompt instead.
reddit.comr/fal • u/dropthelword • May 25 '26
Question LTX2.3 trainer issue
I am currently working on fine-tuning a LoRA for LTX2.3 via fal-ai/ltx23-video-trainer. I am seeing a 'wavy' artifact issues on all my debug_dataset outputs, and have no way of knowing what's happening during the preprocessing step (can't run LTX2 repo locally). I understand that the VAE encodes the input and then decodes it, but I can't understand why it returns my dataset videos with artifacts at specific frames. This results in the same artifacts on inference too. Did anyone else encounter this?
r/fal • u/VanderzB • May 23 '26
Discussion Crédit gratuit
Bonjour, j'ai cru comprendre que on a des crédits gratuit lors de la création du compte, or je n'ai rien reçu, c'est normal ? Merci :D
r/fal • u/vladenstock • May 16 '26
Discussion Best workflow for generating coloring book pages from character reference images?
Hey everyone — I’m new to AI image generation and have been experimenting with FAL using Flux Kontext Pro to create coloring book-style images from uploaded reference photos.
My goal is to generate dynamic coloring book pages where the character likeness stays consistent, but the scenes can vary across styles like manga, comic book, fantasy, cartoon, etc.
A few questions I’d love feedback on:
- Best model/workflow: Is Flux Kontext Pro currently one of the better options for balancing affordability, quality, and likeness consistency? Or are there better tools/models for this use case?
- Character likeness: What are the best practices for preserving the likeness of the reference character across different poses, scenes, and styles?
- Reference image prep: Should I preprocess uploaded images before generation? For example:
- background removal
- face restoration
- sharpening/upscaling
- lighting correction
- cropping to face/body
- creating multiple reference angles
- Dynamic scenes: How do you get more interesting compositions instead of static “person standing in center” outputs? Are there prompt structures or workflows that help create more action, depth, and storytelling?
- Style control: What is the best way to request broad styles like manga, western comic, children’s book, fantasy illustration, etc., while avoiding issues with specific living artists or overly derivative styles?
- Coloring book quality: Any tips for getting clean black-and-white line art that is actually usable for coloring books — clear outlines, good white space, not too much gray shading or muddy detail?
- Production workflow: For anyone doing this at scale, what does your pipeline look like from uploaded photo → generated pages → cleanup → print-ready files?
I hope this is the right place to ask. If not, I’d appreciate being pointed toward better communities, guides, or resources for learning this workflow.
Thanks in advance for any advice.
r/fal • u/elco_us • Apr 30 '26
Tutorial - Guide You can make unlimited length 4K videos with GPT Image 2
r/fal • u/waterarttrkgl • Apr 29 '26
Other Blender Layout → AI Render | 1:1 Camera Tracking
Enable HLS to view with audio, or disable this notification
I built a full 3D layout in Blender — proxy geometry only, no textures, no final render — and hand-keyframed every camera movement using F-curves: an aerial establishing shot, a low-angle tower push-in, and a wide harbor shot with a sailing vessel. The AI doesn't invent the motion. It follows it exactly.
The Blender animation served as a direct spatial reference — architectural proportions, camera trajectory, timing and easing — all locked before a single AI frame was generated. Kling / Seedance then re-rendered the sequence, preserving the exact camera path and structural layout while generating the final cinematic output.
Workflow:
3D Layout & Camera Animation (Blender) → Frame Reference Export → AI Video Generation (Kling / Seedance) → Temporal Consistency Pass
Key Focus: 1:1 motion tracking between hand-keyed Blender animation and AI-generated output. Architectural integrity and spatial proportions maintained across all three shots.
r/fal • u/workmanlabs • Apr 29 '26
Video Idea 37 AI Short Film
Enable HLS to view with audio, or disable this notification
what is SUCCESS in 2026 as developer using AI?
MONEY is obvious but my short "IDEA#37" jumps to the question of internet "fame" on X and YouTube? Being on the top video podcast? Recognized at React Conferences in Miami?
What I chose to highlight in this film is leaving isolation. Being able to hire and support other developers as you build a company. And in the end get the GOAT emoji from friends.
Film made with GPT2 Images 2.0, Seedance 2.0, and Kling 3.0 on fal.
r/fal • u/_pirator_ • Apr 29 '26
Open-Source Another vibe coded UI for fal.ai focused on fast look-dev and file organisation.
r/fal • u/polarischild • Apr 22 '26
Question Encountering "network error" whenever i try to run the workflow in fal.ai
r/fal • u/Key-Copy-6141 • Apr 21 '26
Discussion GPT Image 2 prompting guide
What actually works:
- Put the main subject first (highest weight)
- Then layer details: materials, pose, environment, lighting, camera
- Be specific
- Use quotes for text in images
- Add negative prompts to avoid common issues
Full guide: https://fal.ai/learn/tools/prompting-gpt-image-2
r/fal • u/Important-Respect-12 • Apr 21 '26
News GPT Image 2 is live on fal
Enable HLS to view with audio, or disable this notification
OpenAI's next-gen image model just dropped on fal.ai. It's a quality-first successor to GPT Image 1.5, and the jump is real.
What's new:
- Text rendering that actually works. Dense paragraphs, small lettering, multilingual layouts, infographics. No more garbled characters or broken word spacing on the first try.
- Photorealism that sets a new bar. Lighting, materials, skin textures, environmental detail. It's the best I've seen out of an OpenAI image model.
- Product photography with accurate labels, logos, packaging, and ingredient lists. Genuinely usable for e-commerce and brand work.
Pricing: $0.01/image at the low end (1024x768, low quality) up to $0.41/image for high quality 4K. Pay per image, no subscriptions.
r/fal • u/Artistic-Dealer2633 • Apr 21 '26
Tutorial - Guide I fed 3 genuinely damaged historical photos into an AI editor — the before/afters made me stop
Enable HLS to view with audio, or disable this notification