AI Models & Technology
Powered by the world's strongest AI. Depth Studio unifies the latest
models from industry-leading AI providers under one roof, so a production team can pick the
right engine for each shot without leaving the workspace.
Video generation models
Cinematic quality, high-consistency and professional video synthesis tools.
- Veo — Google's highest quality cinematic video generation architecture.
- Sora — advanced video synthesis with storyboard-driven, multi-shot storytelling.
- Kling — top-tier generation with motion control and lip sync; Kling O3 adds text, frame and reference modes in one model, native audio, 4K and six-shot planning.
- Runway — industry-standard cinematic realism, including video-to-video restyling.
- Seedance — multimodal video generation with lip sync and multi-scene consistency.
- Wan — production-grade video synthesis with fast, cost-efficient modes; Wan 3.0 Video and Wan 3.0 Video Prime accept first and last frames plus up to ten reference images and five reference clips.
- Gemini Omni 1.1 Flash — Google's fast Omni tier: 360p to 4K, four to ten seconds, and editing of a clip you supply.
- Hailuo — creative and dynamic movement generation.
- MiniMax H3 — 2K video from text, first/last frames or reference media, with native audio.
- LTX — 4K 50FPS open-source cinematic video engine.
- Luma Dream Machine — fast, consistent and smooth motion generation.
Image generation models
Visual creation ranging from artistic excellence to technical photorealism.
- Imagen — flawless photorealism with Google's latest architecture.
- Flux — world-class detail, composition and prompt adherence.
- Midjourney — the gold standard for AI aesthetics and style.
- Seedream — ByteDance's high-fidelity image models with strong complex-scene handling.
- Ideogram — high-end typography, graphic design and character consistency.
- Nano Banana — optimised engine for fast, creative visual production.
- Qwen Image — advanced editing and image generation; Qwen Image 3.0 adds 1K/2K output, negative prompts and up to three reference images.
- Grok Imagine Image 2.0 — xAI's image model, from text or from up to five reference images.
Audio and music models
Unique music compositions and realistic voice synthesis technologies.
- Suno — full song generation with music and vocals.
- ElevenLabs — the most realistic AI voices and sound effects, multilingual including Turkish.
- Udio — musical fidelity with unique instrumental depth.
- SFX engine and audio isolation — cinematic sound generation, noise removal and voice separation.
Large language models (LLM)
The most advanced intelligences for complex reasoning, coding and creative writing —
used inside Depth Studio for script analysis, scene breakdown and prompt building.
- Gemini — multi-modal intelligence with a large context window, from the fast Flash tiers up to Gemini 3.1 Pro.
- GPT series — advanced reasoning, planning and multi-step thinking, including the three GPT-5.6 tiers (Luna, Terra, Sol).
- Claude — long document analysis and safe AI processing, up to Claude Opus 5.
- Grok — xAI's chat models (4.3, 4.5, 4.6) with adjustable reasoning.
See how credits and plans work, or read
our production guides.
Depth Studio home ·
About ·
Pricing ·
AI models ·
Contact ·
Resources