March Code
AI video production

Dubler: A Digital Twin That Films Your Reels for You

A SaaS that turns a voice idea into a finished Instagram Reel in about 10 minutes: an AI director writes the script and a storyboard with a cost estimate, and a digital twin with your cloned voice appears on screen instead of you. No camera needed at all.

Client: March Code (our own product)AISaaS
Dubler: A Digital Twin That Films Your Reels for You

29 screens

of a full-fledged SaaS

~10 min

from idea to package

1 photo

for a digital twin

0 shoots

no camera needed at all

01

Challenge

Short vertical videos are the main reach channel for experts and small businesses. But between "I have an idea" and "the video is live" sits a whole production: script, shooting, lighting, editing, subtitles, covers. That's exactly where content plans die: there's no time to shoot, being on camera feels awkward, and an editor is expensive.

We asked ourselves: what if we dropped the shoot altogether? The user provides a photo and five minutes of voice recording once, and from then on their digital double does the on-camera work. All that was left was to build a product that goes from "voice idea" to "ready-to-publish package" in about ten minutes and honestly shows the cost before every generation.

Dubler is the studio's own SaaS product and a showcase at the same time: it's how we demonstrate building commercial AI products end to end, from the LLM pipeline and billing to the affiliate program.

02

Solution

Three steps to a finished video

Share your idea

A voice message, text or a ready-made talking-head clip. Speech is understood word for word, since transcription builds word-level timings

Approve the storyboard

The AI director breaks the speech down by the second, comes up with cutaways, and shows a preview of every scene plus a cost estimate. You're charged only after you say yes

Get the package

A 9:16 video with jump cuts and karaoke-style subtitles, a Stories version without the CTA, three covers and a caption with hashtags, all in one archive

The wow features that make the product

AI director
An editing plan in a single call
  • The LLM builds the whole plan: a hook of ≤5 words, cutaways placed strictly over phrases, pattern interrupts
  • The plan is validated deterministically, so the director can't "break" the video
  • Every cutaway explains why it's there, with its prompt and price
Digital twin
One photo + 5 minutes of voice
  • A lip-synced talking head is generated from a single photo
  • Outfit and location are set in words, like "blue suit" or "neon studio", with no reference images
  • Voice clone with A/B comparison of settings and extra training on problem words
Honest economics
A cost estimate before generation
  • Every generation is priced in advance: the user approves a price, not a surprise at the end
  • Credits: hold → capture in one transaction with a ledger entry, automatic refund on failure
  • The Stories version is built from the same edit: a second format without a second estimate
Assistant
The chat can do everything the buttons can
  • 15 tools, from picking a studio to starting the render, shown as cards right in the conversation
  • Expensive actions go only through a confirmation card with the price
  • The engine is transport-agnostic: web now, Telegram next

The pipeline under the hood

Speech → timings

Transcription with word-level timings; speech boundaries snap to a 30 fps frame grid, so drift against the lip-sync is impossible by design

Director → plan

The LLM returns the editing plan as JSON; the plan passes deterministic validation and becomes a storyboard with a cost estimate

Generation

Video cutaways and voiceover are generated in parallel with retry ×3; BullMQ queues survive restarts, and stale guards finish off stuck jobs

Remotion → package

A 9:16 render with karaoke-style subtitles, a Stories version, three covers and a caption with hashtags, all in one archive

Async as a UX principle

Generation takes minutes, but the user is never blocked: live progress cards with phases and ETAs, skeletons instead of spinners, SSE streams instead of page reloads. The product feels fast even where heavy video models are doing the work under the hood.

03

The system from the inside

Real screens of a working system, not mockups. Click to take a closer look

Product
A showcase of real videos generated by the service, with the owner's face and voice. Not a single frame was shot on camera
The "things to finish" dashboard: videos waiting for a decision, quick actions and the latest finished work, so the user always knows the next step
Video pipeline
The director's storyboard with a cost estimate: every cutaway comes with a preview, prompt, price and the reason it's there. Generation starts only after you say yes
A finished video: the publishing package, edits with a single command ("make it punchier"), a Stories version and the director's report on why it was cut this way
Hybrid cover: the scene is generated with the user's face and the headline is rendered on top, so even Cyrillic text comes out without artifacts
Digital twin
A talking head from one photo: studio locations and wardrobe as cards, or describe them in a single phrase. Then comes animation with lip-sync
Voice clone: up to 5 minutes of recording, voice character picked with chips, A/B comparison of settings and extra training on problem words
Assistant and footage
The assistant chat can do everything the buttons can: studio cards right in the conversation, step-by-step guidance, price confirmation before every charge
Your own footage as the basis for videos: transcription with word search, scene analysis and a "make a video from this footage" button
04

Results

29 screens

of a full-fledged SaaS

Billing with ledger-based credit accounting, an assistant with 15 tools, 34 wow presets, a choice of 3 video models, an affiliate program with a virality calculator

~10 min

from idea to package

A voice idea becomes a 9:16 video with jump cuts and subtitles, a Stories version, three covers and a caption with hashtags, all in one archive

1 photo

for a digital twin

A lip-synced talking head is generated from one photo, and the voice clone from 5 minutes of recording. Outfit and location are set in plain words

0 shoots

no camera needed at all

The product's entire showcase is videos where not a single frame was shot on camera: scenes, cutaways and covers are generated with the owner's face

05

In-depth breakdown

Engineering decisions we're proud of

Money is handled like at a bank
  • Credits: hold → capture/release in one transaction with an append-only ledger
  • Idempotent charges keyed by unique keys, so a double charge can't happen
  • Refunds based on actual usage: if a generation fails, the money comes back on its own
Lip-sync that doesn't break
  • Jump cuts are made on a 30 fps frame grid, so audio and video are cut identically
  • Drift is impossible by design, not just "usually doesn't happen"
Multi-tenancy and security
  • Every business table carries its own orgId; access goes only through a JWT scope
  • Production: Docker Compose, Caddy with auto-TLS, MinIO on a dedicated volume, non-root containers
Resilience to provider outages
  • Transient failures of video models and TTS are retried with backoff
  • Queues survive restarts; an LLM gateway with tiering and caching cuts costs

Want a product like this?

Dubler is how we build commercial AI products end to end: LLM pipelines with deterministic validation, honest ledger-based billing, async UX and production infrastructure. If you have a product idea where AI meets content, bring it to us: we'll build a prototype, work out the unit economics and take it to production. Leave a request below.

Project technologies

NestJS 10TypeORMPostgreSQL 16Redis 7BullMQReact 18ViteTailwindRemotionClaude APIfal.ai (Seedance / Kling)ElevenLabsWhisperMiniMax TTSffmpegDockerMinIO

Want a similar result?

Pilot from $3,900, prototype free

Prices are indicative and not a binding offer.

We build an agent prototype for one scenario using your examples, so you can judge answer quality before signing.

Step 1 of 2 · Task